HTTP API
For AI agents: see llms.txt for the complete documentation index. Markdown versions are available by adding .md to a page URL or requesting Accept: text/markdown.
Base URL: https://ai-gateway.convex.dev. To send a request from an action, see
Getting started.
Authentication
Authorization: Bearer <token>
Get the token in an action with getServiceToken("ai-gateway"). The action
runtime caches and refreshes it as needed. Token rules and errors:
Getting started.
Missing or invalid token:
{
"error": {
"message": "Invalid authentication credentials",
"type": "invalid_request_error",
"code": "invalid_api_key"
}
}
Provider preferences
Chat Completions, Embeddings, Messages, Responses, and Decisions accept these
optional fields inside provider:
| Field | Accepted values | Purpose |
|---|---|---|
zdr | Boolean | Require zero-data-retention inference endpoints when true. |
data_collection | "deny" | Exclude providers that collect data. |
require_parameters | Boolean | Require providers to support all supplied parameters when true. |
max_price | Object | Limit acceptable provider rates. |
{
"model": "openai/gpt-4o-mini",
"messages": [{ "role": "user", "content": "Hello!" }],
"provider": {
"zdr": true,
"data_collection": "deny",
"require_parameters": true,
"max_price": { "prompt": "1", "completion": "2" }
}
}
max_price accepts prompt and completion (USD per million tokens),
request (USD per request), image (USD per image), and audio (USD per audio
unit). Values must be finite, nonnegative numbers or numeric strings. These cap
provider prices, not total request cost; see
usage limits.
Convex rejects other provider fields and data_collection values other than
"deny" with HTTP 400. The upstream service validates the remaining values.
Omitting provider or sending {} uses the default policy. Image, video, and
audio routes reject provider entirely.
zdr: true restricts inference routing to endpoints with a zero-data-retention
policy. See Models for
catalog availability. If no eligible endpoint matches your request, the request
fails. zdr: false cannot disable a stricter gateway policy.
ZDR routing does not govern your application's logging or third-party tools. It may allow provider-side in-memory prompt caching. Usage and billing metadata are still retained.
GET /v1/models
No query parameters or body.
{
"object": "list",
"data": [
{
"id": "openai/gpt-4o-mini",
"object": "model",
"created": 1715620800,
"owned_by": "openai"
}
]
}
owned_by is the provider prefix of id. The list includes text, embedding,
image, video, speech, transcription, and Decisions models. It also includes
rerank models, which the gateway does not serve.
POST /v1/chat/completions
OpenAI Chat Completions
body. Set stream: true for
SSE.
| Field | Type | Required | Description |
|---|---|---|---|
| model | string | y | provider/model id |
| messages | array | y | OpenAI messages |
| stream | boolean | n | SSE when true. Defaults to false |
Other OpenAI fields (temperature, max_tokens, tools, response_format, …)
are forwarded. Body must be JSON, max 16 MiB.
The provider preferences above are supported. These
fields are rejected: route, models, transforms, plugins, preset.
{
"error": {
"message": "The `provider.only` parameter is invalid or unsupported for this request.",
"type": "invalid_request_error",
"code": "unsupported_parameter",
"param": "provider.only"
}
}
Response
id is assigned by Convex. Non-streaming is application/json:
{
"id": "3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90",
"object": "chat.completion",
"created": 1715367049,
"model": "openai/gpt-4o-mini",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello!" },
"finish_reason": "stop",
"logprobs": null
}
],
"system_fingerprint": "fp_123",
"usage": {
"prompt_tokens": 12,
"completion_tokens": 5,
"total_tokens": 17,
"prompt_tokens_details": { "cached_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 }
}
}
Streaming is text/event-stream. Chunks use "object": "chat.completion.chunk"
and choices[].delta instead of choices[].message:
data: {"id":"3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90","object":"chat.completion.chunk","created":1715367049,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: [DONE]
Retry-After is forwarded when present.
POST /v1/embeddings
OpenAI Embeddings
body. model and input are required. input may be one string, one token-ID
array, or a batch of up to 512 strings or token-ID arrays. Other OpenAI fields
are forwarded. The body must be JSON and no larger than 16 MiB.
Convex rejects the same routing controls as /v1/chat/completions, assigns the
response id, and removes fields that identify the serving provider. The
response otherwise uses the OpenAI embeddings shape.
POST /v1/images/generations
The request and response format may change during alpha.
Send a JSON body with model and prompt, plus model-supported image options
such as size and n. Each data entry in the response contains b64_json
and, when known, media_type. Streaming is rejected. The same routing controls
as /v1/chat/completions, plus provider, are rejected. The body limit is 16
MiB.
POST /v1/videos/generations
The request and response format may change during alpha.
Send model and prompt, with optional duration, aspect_ratio, size,
resolution, seed, generate_audio, frame_images, or input_references.
Supported values depend on the model. Each request generates one video; n,
stream, callback_url, and provider are rejected, along with gateway
routing controls. The body limit is 16 MiB.
This request waits for completion and downloads the video, with a 64 MiB video
limit. It fails if generation takes longer than about five minutes. Use the
async video routes for longer jobs. The JSON response
contains a Convex-assigned id, data: [{ b64_json, media_type }], and usage
when available. Cancelling the HTTP request does not cancel generation or its
cost.
POST /v1/audio/transcriptions
The request and response format may change during alpha.
Send a JSON body with model and input_audio: { data, format }. data is
base64 audio with no data: prefix. format names the encoding, such as wav,
mp3, flac, m4a, ogg, webm, or aac. Supported formats vary by
provider. Optional fields are language, temperature, response_format, and
timestamp_granularities. The gateway rejects multipart uploads with
unsupported_media_type. It also rejects stream: true and the routing
controls it rejects on /v1/chat/completions. The body limit is 16 MiB.
The response contains text and usage with seconds, token counts when the
model reports them, and cost. Set response_format: "verbose_json" to request
structured fields such as task, language, duration, and segments.
Supported fields depend on the provider and model. Add
timestamp_granularities: ["word"] to request words. Models that do not
support verbose_json return a 400 error.
POST /v1/audio/speech
The request and response format may change during alpha.
OpenAI Speech
body. model and input are required. Set voice unless the selected provider
documents a default voice. response_format (mp3 or pcm) and speed are
optional. response_format defaults to pcm. The gateway rejects the routing
controls it rejects on /v1/chat/completions.
A successful response is the raw audio. Its Content-Type is audio/pcm for
pcm (16-bit little-endian) and audio/mpeg for mp3. Errors are JSON.
Async video routes
These routes are in alpha. For availability and local development, see
Async videos. All three routes
require a deployment token. Save the opaque operation value returned at
submission and use a fresh token from the same deployment for later requests.
Operations expire after seven days; video retention may be shorter.
POST /v1/videos
Accepts the same generation options as /v1/videos/generations, plus an
optional webhook_url on the deployment's HTTPS <deployment>.convex.site
origin. Returns HTTP 202:
{
"id": "convex-inference-id",
"operation": "opaque-operation-handle",
"webhook_secret": "per-job-signing-secret"
}
webhook_secret is null when webhook_url is omitted. Store it privately. A
lost response can still mean the job was accepted and charged; automatically
retrying submission can create another paid job.
POST /v1/videos/status
Send { "operation": "opaque-operation-handle" }. Returns status as
pending, completed, or error, with an error message for failed jobs and
usage when available.
POST /v1/videos/download
Send the same operation body. Returns the video in the same data shape as
/v1/videos/generations. Returns 409 if the video is not ready. Each call
fetches the video again; save it in application storage.
Application callbacks
The gateway posts a signed JSON event to webhook_url with id, operation,
status, and type (video.generation.<status>). Terminal statuses are
completed, failed, cancelled, and expired.
Verify the raw body and x-convex-video-signature header with
verifyVideoWebhook before changing application state. Match the event's id
to the saved inference ID, and deduplicate by (id, status) when saving it. See
receiving a callback.
Respond within 8 seconds, or the delivery fails. Callbacks can repeat, and a
failed delivery may not be retried, so check unfinished jobs periodically.
POST /alpha/decisions
The request and response format may change during alpha.
Jev is TypeSafe's model for making structured decisions: choosing an option, scoring data, or evaluating a statement.
Provide the context to evaluate in state and the questions to answer in
questions. You can ask several questions about the same state in one request.
Each question is evaluated independently and has a name that identifies its
answer in the response.
model, state, and questions are required. Use typesafe/jev-1.13 as the
model ID. Streaming is not supported.
With the AI SDK provider, use
evaluate({ model: convexGateway.evaluationModel("typesafe/jev-1.13"), ... }).
The SDK calls the yes-or-no question type boolean and returns probability;
the HTTP API uses noul for both.
For example, classify a support ticket by priority:
{
"model": "typesafe/jev-1.13",
"state": {
"ticket": "All users are unable to sign in. There is no workaround."
},
"questions": {
"priority": {
"type": "choice",
"instructions": "Choose the support ticket's priority.",
"criteria": {
"urgent": "An outage is blocking users.",
"normal": "A bug affects users but has a workaround."
}
}
}
}
state and each question's instructions accept a string, object, or array.
Each question must specify one of these types:
type | Purpose | Required fields | Result field |
|---|---|---|---|
choice | Select an option | instructions, criteria map | choice |
score | Evaluate the state against a rubric | instructions, criteria array | score |
noul | Evaluate whether a statement is true | instructions | noul, a number from 0–1 |
Keep each question focused on one decision. For decisions involving several factors, ask about each factor separately and combine the answers in your code.
A noul question may include criteria with true and false guidance.
Choice criteria values may be strings, objects, arrays, or null. Score
criteria are an ordered array of strings, objects, or arrays.
The body must be JSON and no larger than 16 MiB. These fields are not supported:
route, models, transforms, plugins, preset, fallbacks, speed,
trace, session_id, user, and stream.
Response
The answer keys match the request's question keys. Convex assigns id and
removes fields that identify the serving provider.
choice and score answers may also include confidence and probabilities.
{
"id": "3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90",
"model": "typesafe/jev-1.13",
"answers": {
"priority": {
"type": "choice",
"choice": "urgent",
"confidence": 0.9,
"probabilities": { "urgent": 0.9, "normal": 0.1 }
}
},
"usage": {
"input_tokens": 21,
"output_tokens": 3,
"cost": 0.0042
}
}
usage contains the input and output token counts. usage.cost, when present,
is the request cost in US dollars.
POST /v1/messages
Anthropic Messages body. Model IDs
use the provider/model form. model, messages, and max_tokens are
required. Set stream: true for Anthropic-compatible server-sent events.
{
"model": "anthropic/claude-haiku-4.5",
"max_tokens": 128,
"messages": [{ "role": "user", "content": "Hello!" }]
}
Other Anthropic fields are forwarded. The body must be JSON and no larger than
16 MiB. These routing controls are rejected because Convex chooses how the
request is served: route, models, plugins, fallbacks, session_id, and
speed.
The response uses the Anthropic Messages shape. Convex replaces upstream id
and request_id values with Convex-generated IDs and removes fields that
identify the serving provider. Local errors, including authentication errors,
also use the Anthropic error shape:
{
"type": "error",
"error": {
"type": "authentication_error",
"message": "Invalid authentication credentials"
},
"request_id": "3f1c8a2e-9b14-4d6a-a7e2-0c5b8d1e4f90"
}
POST /v1/responses
OpenAI Responses
body. Model IDs use the provider/model form. model and input are required.
Set stream: true for server-sent events.
{
"model": "openai/gpt-5-mini",
"input": "Hello!"
}
Other OpenAI Responses fields are forwarded. The body must be JSON and no larger
than 16 MiB. These routing controls are rejected because Convex chooses how the
request is served: route, models, transforms, plugins, preset, and
session_id.
The endpoint is stateless. The upstream service rejects store: true and a
non-null previous_response_id. Convex replaces upstream response IDs with
Convex-generated IDs and removes fields that identify the serving provider.
Caller-supplied metadata on a response is preserved.
Errors
| Status | code | When |
|---|---|---|
| 401 | invalid_api_key | Missing or invalid Authorization |
| 400 | unsupported_endpoint | Unknown path, with a valid token |
| 400 | unsupported_parameter | Rejected routing field |
| 400 | too_many_inputs | Over 512 embedding inputs |
| 413 | request_too_large | Body over 16 MiB |
| 415 | unsupported_media_type | Multipart transcription upload |
| 502 / 503 | upstream_error | Provider temporarily unavailable |
| 502 / 504 | video_generation_failed | Synchronous video generation failed or timed out |
Provider validation errors (unknown model, bad args) keep the provider status.
The error object has message, type, code, and param when present.