Skip to main content

Voice

For AI agents: see llms.txt for the complete documentation index. Markdown versions are available by adding .md to a page URL or requesting Accept: text/markdown.

Turn audio into text and text into audio from an action with convexGateway. For installation and authentication, see Getting started.

Voice needs @convex-dev/ai-sdk-provider 0.3.0-alpha.0 or later:

npm install @convex-dev/ai-sdk-provider@alpha
Voice is in alpha

The request and response format may change during alpha.

The gateway supports three voice modes:

ModeIn → outUse it to
Speech-to-textAudio → textTranscribe recordings, voice notes, meetings
Text-to-speechText → audioRead text aloud in a chosen voice
Audio chatAudio or text → text, audioLet a model hear a question and answer it

To have a model answer by voice, chain transcription and speech with a chat model, or use an audio chat model.

Send each recording in a separate request. The gateway can stream audio replies. It does not support two-way live audio sessions. Your app controls playback and handles interruptions.

The examples below run inside an action handler.

Transcribe audio​

storageId is the ID of an audio file in file storage.

import { experimental_transcribe as transcribe } from "ai";
import { convexGateway } from "@convex-dev/ai-sdk-provider";

const { text } = await transcribe({
model: convexGateway.transcriptionModel("openai/whisper-large-v3-turbo"),
audio: await (await ctx.storage.get(storageId))!.arrayBuffer(),
});

Set the language with providerOptions: { convexGateway: { language: "en" } }. The model detects it when you leave it out.

The request body limit is 16 MiB. Base64 encoding makes audio a third larger, so one request holds about 12 MiB of raw audio. Providers also time out after 60 seconds of processing. Split longer recordings into shorter clips.

Generate speech​

import { experimental_generateSpeech as generateSpeech } from "ai";
import { convexGateway } from "@convex-dev/ai-sdk-provider";

const { audio } = await generateSpeech({
model: convexGateway.speechModel("hexgrad/kokoro-82m"),
text: "Your order has shipped.",
voice: "af_heart",
});

const storageId = await ctx.storage.store(
new Blob([audio.uint8Array], { type: audio.mediaType }),
);

Voices depend on the model. Always pass a voice the model supports. Without one, the SDK sends OpenAI's alloy, which most other models do not offer. The response is MP3. Set outputFormat: "pcm" for raw PCM. The SDK still labels PCM audio audio/mp3, so set the Blob type to audio/pcm yourself.

Answer by voice​

Chain the three steps to answer a spoken question with any chat model:

import {
experimental_generateSpeech as generateSpeech,
experimental_transcribe as transcribe,
generateText,
} from "ai";
import { convexGateway } from "@convex-dev/ai-sdk-provider";

const { text: question } = await transcribe({
model: convexGateway.transcriptionModel("openai/whisper-large-v3-turbo"),
audio: recording, // ArrayBuffer of the caller's audio
});

const { text: answer } = await generateText({
model: convexGateway("anthropic/claude-sonnet-4.5"),
prompt: question,
});

const { audio } = await generateSpeech({
model: convexGateway.speechModel("hexgrad/kokoro-82m"),
text: answer,
voice: "af_heart",
});

An audio chat model such as openai/gpt-audio does all three in one call through /v1/chat/completions. Send audio as an input_audio content part. For spoken output, set modalities: ["text", "audio"] and an audio object with voice and format, and stream the response. The audio arrives in delta.audio chunks. Use the OpenAI SDK or fetch for spoken output, because the AI SDK returns only the text.

Models​

Start with these models. Call GET /v1/models for every model the gateway serves.

ModeModels
Speech-to-textopenai/whisper-large-v3-turbo, openai/gpt-4o-transcribe, deepgram/nova-3
Text-to-speechhexgrad/kokoro-82m, google/gemini-3.8-flash-tts, deepgram/aura-2
Audio chatopenai/gpt-audio, openai/gpt-audio-mini

Cost​

Transcription results include their cost in providerMetadata.convexGateway.cost, in US dollars. Speech results have no cost field, because the cost is known only after the audio is sent. Convex still bills for speech. The charge appears on the team usage page.