Skip to main content

HTTP API

For AI agents: see llms.txt for the complete documentation index. Markdown versions are available by adding .md to a page URL or requesting Accept: text/markdown.

Base URL: https://ai-gateway.convex.dev. To send a request from an action, see Getting started.

Authentication​

Authorization: Bearer <token>

Get the token in an action with getServiceToken("ai-gateway"). The action runtime caches and refreshes it as needed. Token rules and errors: Getting started.

Missing or invalid token:

{
"error": {
"message": "Invalid authentication credentials",
"type": "invalid_request_error",
"code": "invalid_api_key"
}
}

Provider preferences​

Chat Completions, Embeddings, Messages, Responses, and Decisions accept these optional fields inside provider:

FieldAccepted valuesPurpose
zdrBooleanRequire zero-data-retention inference endpoints when true.
data_collection"deny"Exclude providers that collect data.
require_parametersBooleanRequire providers to support all supplied parameters when true.
max_priceObjectLimit acceptable provider rates.
{
"model": "openai/gpt-4o-mini",
"messages": [{ "role": "user", "content": "Hello!" }],
"provider": {
"zdr": true,
"data_collection": "deny",
"require_parameters": true,
"max_price": { "prompt": "1", "completion": "2" }
}
}

max_price accepts prompt and completion (USD per million tokens), request (USD per request), image (USD per image), and audio (USD per audio unit). Values must be finite, nonnegative numbers or numeric strings. These cap provider prices, not total request cost; see usage limits.

Convex rejects other provider fields and data_collection values other than "deny" with HTTP 400. The upstream service validates the remaining values. Omitting provider or sending {} uses the default policy. Image, video, and audio routes reject provider entirely.

zdr: true restricts inference routing to endpoints with a zero-data-retention policy. See Models for catalog availability. If no eligible endpoint matches your request, the request fails. zdr: false cannot disable a stricter gateway policy.

ZDR routing does not govern your application's logging or third-party tools. It may allow provider-side in-memory prompt caching. Usage and billing metadata are still retained.

GET /v1/models​

No query parameters or body.

{
"object": "list",
"data": [
{
"id": "openai/gpt-4o-mini",
"object": "model",
"created": 1715620800,
"owned_by": "openai"
}
]
}

owned_by is the provider prefix of id. The list includes text, embedding, image, video, speech, transcription, and Decisions models. It also includes rerank models, which the gateway does not serve.

POST /v1/chat/completions​

OpenAI Chat Completions body. Set stream: true for SSE.

FieldTypeRequiredDescription
modelstringyprovider/model id
messagesarrayyOpenAI messages
streambooleannSSE when true. Defaults to false

Other OpenAI fields (temperature, max_tokens, tools, response_format, …) are forwarded. Body must be JSON, max 16 MiB.

The provider preferences above are supported. These fields are rejected: route, models, transforms, plugins, preset.

{
"error": {
"message": "The `provider.only` parameter is invalid or unsupported for this request.",
"type": "invalid_request_error",
"code": "unsupported_parameter",
"param": "provider.only"
}
}

Response​

id is assigned by Convex. Non-streaming is application/json:

{
"id": "3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90",
"object": "chat.completion",
"created": 1715367049,
"model": "openai/gpt-4o-mini",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello!" },
"finish_reason": "stop",
"logprobs": null
}
],
"system_fingerprint": "fp_123",
"usage": {
"prompt_tokens": 12,
"completion_tokens": 5,
"total_tokens": 17,
"prompt_tokens_details": { "cached_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 }
}
}

Streaming is text/event-stream. Chunks use "object": "chat.completion.chunk" and choices[].delta instead of choices[].message:

data: {"id":"3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90","object":"chat.completion.chunk","created":1715367049,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: [DONE]

Retry-After is forwarded when present.

POST /v1/embeddings​

OpenAI Embeddings body. model and input are required. input may be one string, one token-ID array, or a batch of up to 512 strings or token-ID arrays. Other OpenAI fields are forwarded. The body must be JSON and no larger than 16 MiB.

Convex rejects the same routing controls as /v1/chat/completions, assigns the response id, and removes fields that identify the serving provider. The response otherwise uses the OpenAI embeddings shape.

POST /v1/images/generations​

Image generation is in alpha

The request and response format may change during alpha.

Send a JSON body with model and prompt, plus model-supported image options such as size and n. Each data entry in the response contains b64_json and, when known, media_type. Streaming is rejected. The same routing controls as /v1/chat/completions, plus provider, are rejected. The body limit is 16 MiB.

POST /v1/videos/generations​

Video generation is in alpha

The request and response format may change during alpha.

Send model and prompt, with optional duration, aspect_ratio, size, resolution, seed, generate_audio, frame_images, or input_references. Supported values depend on the model. Each request generates one video; n, stream, callback_url, and provider are rejected, along with gateway routing controls. The body limit is 16 MiB.

This request waits for completion and downloads the video, with a 64 MiB video limit. It fails if generation takes longer than about five minutes. Use the async video routes for longer jobs. The JSON response contains a Convex-assigned id, data: [{ b64_json, media_type }], and usage when available. Cancelling the HTTP request does not cancel generation or its cost.

POST /v1/audio/transcriptions​

Voice is in alpha

The request and response format may change during alpha.

Send a JSON body with model and input_audio: { data, format }. data is base64 audio with no data: prefix. format names the encoding, such as wav, mp3, flac, m4a, ogg, webm, or aac. Supported formats vary by provider. Optional fields are language, temperature, response_format, and timestamp_granularities. The gateway rejects multipart uploads with unsupported_media_type. It also rejects stream: true and the routing controls it rejects on /v1/chat/completions. The body limit is 16 MiB.

The response contains text and usage with seconds, token counts when the model reports them, and cost. Set response_format: "verbose_json" to request structured fields such as task, language, duration, and segments. Supported fields depend on the provider and model. Add timestamp_granularities: ["word"] to request words. Models that do not support verbose_json return a 400 error.

POST /v1/audio/speech​

Voice is in alpha

The request and response format may change during alpha.

OpenAI Speech body. model and input are required. Set voice unless the selected provider documents a default voice. response_format (mp3 or pcm) and speed are optional. response_format defaults to pcm. The gateway rejects the routing controls it rejects on /v1/chat/completions.

A successful response is the raw audio. Its Content-Type is audio/pcm for pcm (16-bit little-endian) and audio/mpeg for mp3. Errors are JSON.

Async video routes​

These routes are in alpha. For availability and local development, see Async videos. All three routes require a deployment token. Save the opaque operation value returned at submission and use a fresh token from the same deployment for later requests. Operations expire after seven days; video retention may be shorter.

POST /v1/videos​

Accepts the same generation options as /v1/videos/generations, plus an optional webhook_url on the deployment's HTTPS <deployment>.convex.site origin. Returns HTTP 202:

{
"id": "convex-inference-id",
"operation": "opaque-operation-handle",
"webhook_secret": "per-job-signing-secret"
}

webhook_secret is null when webhook_url is omitted. Store it privately. A lost response can still mean the job was accepted and charged; automatically retrying submission can create another paid job.

POST /v1/videos/status​

Send { "operation": "opaque-operation-handle" }. Returns status as pending, completed, or error, with an error message for failed jobs and usage when available.

POST /v1/videos/download​

Send the same operation body. Returns the video in the same data shape as /v1/videos/generations. Returns 409 if the video is not ready. Each call fetches the video again; save it in application storage.

Application callbacks​

The gateway posts a signed JSON event to webhook_url with id, operation, status, and type (video.generation.<status>). Terminal statuses are completed, failed, cancelled, and expired.

Verify the raw body and x-convex-video-signature header with verifyVideoWebhook before changing application state. Match the event's id to the saved inference ID, and deduplicate by (id, status) when saving it. See receiving a callback. Respond within 8 seconds, or the delivery fails. Callbacks can repeat, and a failed delivery may not be retried, so check unfinished jobs periodically.

POST /alpha/decisions​

Decisions is in alpha

The request and response format may change during alpha.

Jev is TypeSafe's model for making structured decisions: choosing an option, scoring data, or evaluating a statement.

Provide the context to evaluate in state and the questions to answer in questions. You can ask several questions about the same state in one request. Each question is evaluated independently and has a name that identifies its answer in the response.

model, state, and questions are required. Use typesafe/jev-1.13 as the model ID. Streaming is not supported.

With the AI SDK provider, use evaluate({ model: convexGateway.evaluationModel("typesafe/jev-1.13"), ... }). The SDK calls the yes-or-no question type boolean and returns probability; the HTTP API uses noul for both.

For example, classify a support ticket by priority:

{
"model": "typesafe/jev-1.13",
"state": {
"ticket": "All users are unable to sign in. There is no workaround."
},
"questions": {
"priority": {
"type": "choice",
"instructions": "Choose the support ticket's priority.",
"criteria": {
"urgent": "An outage is blocking users.",
"normal": "A bug affects users but has a workaround."
}
}
}
}

state and each question's instructions accept a string, object, or array. Each question must specify one of these types:

typePurposeRequired fieldsResult field
choiceSelect an optioninstructions, criteria mapchoice
scoreEvaluate the state against a rubricinstructions, criteria arrayscore
noulEvaluate whether a statement is trueinstructionsnoul, a number from 0–1

Keep each question focused on one decision. For decisions involving several factors, ask about each factor separately and combine the answers in your code.

A noul question may include criteria with true and false guidance. Choice criteria values may be strings, objects, arrays, or null. Score criteria are an ordered array of strings, objects, or arrays.

The body must be JSON and no larger than 16 MiB. These fields are not supported: route, models, transforms, plugins, preset, fallbacks, speed, trace, session_id, user, and stream.

Response​

The answer keys match the request's question keys. Convex assigns id and removes fields that identify the serving provider.

choice and score answers may also include confidence and probabilities.

{
"id": "3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90",
"model": "typesafe/jev-1.13",
"answers": {
"priority": {
"type": "choice",
"choice": "urgent",
"confidence": 0.9,
"probabilities": { "urgent": 0.9, "normal": 0.1 }
}
},
"usage": {
"input_tokens": 21,
"output_tokens": 3,
"cost": 0.0042
}
}

usage contains the input and output token counts. usage.cost, when present, is the request cost in US dollars.

POST /v1/messages​

Anthropic Messages body. Model IDs use the provider/model form. model, messages, and max_tokens are required. Set stream: true for Anthropic-compatible server-sent events.

{
"model": "anthropic/claude-haiku-4.5",
"max_tokens": 128,
"messages": [{ "role": "user", "content": "Hello!" }]
}

Other Anthropic fields are forwarded. The body must be JSON and no larger than 16 MiB. These routing controls are rejected because Convex chooses how the request is served: route, models, plugins, fallbacks, session_id, and speed.

The response uses the Anthropic Messages shape. Convex replaces upstream id and request_id values with Convex-generated IDs and removes fields that identify the serving provider. Local errors, including authentication errors, also use the Anthropic error shape:

{
"type": "error",
"error": {
"type": "authentication_error",
"message": "Invalid authentication credentials"
},
"request_id": "3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90"
}

POST /v1/responses​

OpenAI Responses body. Model IDs use the provider/model form. model and input are required. Set stream: true for server-sent events.

{
"model": "openai/gpt-5-mini",
"input": "Hello!"
}

Other OpenAI Responses fields are forwarded. The body must be JSON and no larger than 16 MiB. These routing controls are rejected because Convex chooses how the request is served: route, models, transforms, plugins, preset, and session_id.

The endpoint is stateless. The upstream service rejects store: true and a non-null previous_response_id. Convex replaces upstream response IDs with Convex-generated IDs and removes fields that identify the serving provider. Caller-supplied metadata on a response is preserved.

Errors​

StatuscodeWhen
401invalid_api_keyMissing or invalid Authorization
400unsupported_endpointUnknown path, with a valid token
400unsupported_parameterRejected routing field
400too_many_inputsOver 512 embedding inputs
413request_too_largeBody over 16 MiB
415unsupported_media_typeMultipart transcription upload
502 / 503upstream_errorProvider temporarily unavailable
502 / 504video_generation_failedSynchronous video generation failed or timed out

Provider validation errors (unknown model, bad args) keep the provider status. The error object has message, type, code, and param when present.