AI Gateway now handles audio alongside text, image, and video. Realtime voice, text to speech, and speech to text are available through the same calls, with the same provider routing, observability, spend controls, and bring-your-own-key support. Launch models come from OpenAI and xAI. The three capabilities are in beta on AI SDK 7.

Capability

How it works

Use it for

Realtime voice

Live audio in and out, for streaming, low-latency session

Two-way voice agents and live conversation

Text to speech

Text in, audio file out, single request

Voiceovers, spoken responses, audio versions of written content

Speech to text

Recorded audio in, text out, single request

Transcribing voice notes, call recordings

Realtime voice agents

A realtime model is not a chain. Rather than stitching together speech-to-text, a language model, and text-to-speech, one model takes in audio and emits audio, so it can answer the moment a user speaks. That immediacy is what makes barge-in work: the user talks over the reply to cut it short, the way they would with another person. Voice assistants, support agents, and hands-free tools are the obvious fits.

Inside a session, two behaviors differ from an ordinary model call:

  • turnDetection: { type: 'server-vad' } hands turn boundaries to the server, which decides when the user has stopped speaking and supports barge-in — no client-side silence timers.

  • Tool calls can arrive mid-reply. You execute the call and return the result as a client event; the model weaves it into the rest of its response instead of closing the turn.

On AI SDK 7, the browser-side useRealtime hook owns the WebSocket connection, microphone capture, and audio playback. Authentication runs through your AI Gateway credential: mint a short-lived token on the server and send only that token to the client, so the API key stays server-side. Start with a token route, then connect from a client component.

npm install ai @ai-sdk/react @ai-sdk/gateway

import { gateway } from '@ai-sdk/gateway';

export async function POST() {

const { token, url } = await gateway.experimental_realtime.getToken({

model: 'openai/gpt-realtime-2',

});

return Response.json({ token, url, tools: [] });

}

That path captures the mic, streams audio to the model via AI Gateway, and plays the spoken response. For non-browser clients, drive the session over a WebSocket with getWebSocketConfig, serializeClientEvent, and parseServerEvent, as documented in the realtime reference.

Speech in both directions

generateSpeech turns text into spoken audio given a voice and an output format, which you then write to a file.

'use client';

import { experimental_useRealtime as useRealtime } from '@ai-sdk/react';

import { gateway } from '@ai-sdk/gateway';

import { useMemo } from 'react';

export default function Page() {

const model = useMemo(

() => gateway.experimental_realtime('openai/gpt-realtime-2'),

[],

);

const { status, connect, startAudioCapture } = useRealtime({

model,

api: { token: '/api/realtime/token' },

sessionConfig: { voice: 'alloy', turnDetection: { type: 'server-vad' } },

});

// Call connect(), then startAudioCapture(stream) with a microphone MediaStream.

}

transcribe goes the other way, accepting a buffer, a base64 string, or a URL.

import { generateSpeech } from 'ai';

import { writeFile } from 'node:fs/promises';

const result = await generateSpeech({

model: 'xai/grok-tts',

text: 'Thanks for trying out AI Gateway.',

voice: 'eve',

outputFormat: 'mp3',

});

await writeFile('speech.mp3', result.audio.uint8Array);

Since the two are complementary, they compose into a single pass: generate audio with one model, read it back with the other. That is a fast sanity check on both ends of an audio pipeline.

Same routing, same controls

Audio requests follow the pattern already used elsewhere on AI Gateway. One API key covers multiple providers; traffic and usage appear in observability; the same budgets and spend limits apply; and bringing your own provider keys remains an option. Apps already routing text, images, or video through the gateway can add speech in the same place.

Reference material: the realtime quickstart, the speech quickstart for text to speech and speech to text, the realtime reference, and the full audio model list.

Trying models without code

The models page also serves as a browser playground. Open a model and interact with it directly: hold a voice conversation with a realtime model, or send text or audio to a speech or transcription model and read or play back what comes out.