1# Realtime and audio1# Getting started with the Realtime API
2 2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4 4
5Start with the outcome you want to build. Realtime sessions are best for live audio that needs low latency. Request-based audio APIs are best for files, bounded requests, or generated speech that doesn't need a live session.5Build a speech-to-speech voice agent with the Realtime API. The model works directly with audio, maintains conversation state, and can call tools. This guide starts with the Agents SDK for a browser application; use the lower-level connection guides when you need direct control.
6
7## Common use cases
8
9
10
11 - **[Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)**: Build speech-to-speech agents that listen, reason, speak, and call tools.
12- **[Live translation](https://developers.openai.com/api/docs/guides/realtime-translation)**: Translate live speech with a dedicated realtime translation session.
13- **[Transcription](https://developers.openai.com/api/docs/guides/transcription)**: Stream live transcript deltas or process audio files into text.
14- **[Speech generation](https://developers.openai.com/api/docs/guides/text-to-speech)**: Turn text into natural-sounding spoken audio.
15
16
17
18## Understand different architectures
19
20<table>
21 <thead>
22 <tr>
23 <th>Goal</th>
24 <th>Model or API</th>
25 <th>Start here</th>
26 </tr>
27 </thead>
28 <tbody>
29 <tr>
30 <td>Build a low-latency voice agent</td>
31 <td className="whitespace-nowrap">
32 [`gpt-realtime-2.1`](https://developers.openai.com/api/docs/models/gpt-realtime-2.1)
33 </td>
34 <td>
35 [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)
36 </td>
37 </tr>
38 <tr>
39 <td>Translate live speech into another language</td>
40 <td className="whitespace-nowrap">
41 [`gpt-realtime-translate`](https://developers.openai.com/api/docs/models/gpt-realtime-translate)
42 </td>
43 <td>
44 [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation)
45 </td>
46 </tr>
47 <tr>
48 <td>Transcribe live audio into streaming text</td>
49 <td className="whitespace-nowrap">
50 [`gpt-live-transcribe`](https://developers.openai.com/api/docs/models/gpt-live-transcribe)
51 </td>
52 <td>
53 [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)
54 </td>
55 </tr>
56 <tr>
57 <td>Transcribe files or bounded audio requests</td>
58 <td>Audio transcription models</td>
59 <td>
60 [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text)
61 </td>
62 </tr>
63 <tr>
64 <td>Generate speech from text</td>
65 <td>Speech generation models</td>
66 <td>
67 [Text to speech](https://developers.openai.com/api/docs/guides/text-to-speech)
68 </td>
69 </tr>
70 <tr>
71 <td>Add audio to an existing Chat Completions app</td>
72 <td>Audio-capable chat models</td>
73 <td>
74 [Audio and speech](https://developers.openai.com/api/docs/guides/audio#add-audio-to-your-existing-application)
75 </td>
76 </tr>
77 </tbody>
78</table>
79
80## Choose a realtime session
81
82Realtime sessions keep a connection open while your application sends audio, receives events, and updates session state.
83
84<table>
85 <thead>
86 <tr>
87 <th>Session type</th>
88 <th>Use when</th>
89 <th>Endpoint or pattern</th>
90 </tr>
91 </thead>
92 <tbody>
93 <tr>
94 <td>Voice-agent session</td>
95 <td>
96 The model should respond to the user, call tools, and manage
97 conversation state.
98 </td>
99 <td>
100 Conversation session on `/v1/realtime`
101 </td>
102 </tr>
103 <tr>
104 <td>Translation session</td>
105 <td>The app should continuously translate speech as it arrives.</td>
106 <td>
107 Continuous translation session on `/v1/realtime/translations`
108 </td>
109 </tr>
110 <tr>
111 <td>Transcription session</td>
112 <td>
113 The app needs streaming transcript deltas without model-generated spoken
114 responses.
115 </td>
116 <td>Transcription session that emits transcript deltas</td>
117 </tr>
118 </tbody>
119</table>
120
121Use a voice-agent session when your application needs an assistant that responds to the user. Use a translation session when your application needs an interpreter that translates the speaker. Use a transcription session when your application needs text from audio without model-generated responses.
122
123### Voice-agent sessions
124
125Voice-agent sessions use the standard Realtime API conversation lifecycle. The client connects to `/v1/realtime`, sends audio or text, and listens for model responses, tool calls, and session events.
126
127For most browser voice agents, start with the [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) guide. It uses the Agents SDK with WebRTC for browser audio and can connect to server-side tools.
128
129Realtime 2 adds reasoning to speech-to-speech workflows. Start with
130 `reasoning.effort` set to `low` for most production voice agents, then adjust
131 based on latency tolerance and task complexity. Use the [Realtime prompting
132 guide](https://developers.openai.com/api/docs/guides/realtime-models-prompting) to tune reasoning,
133 preambles, tool use, unclear audio, and exact entity capture.
134
135### Translation sessions
136
137Realtime translation uses a dedicated translation endpoint instead of the standard voice-agent endpoint. Translation sessions are continuous: the client streams audio into the session, and the service streams translated audio and transcript deltas out.
138
139Translation sessions don't use the normal assistant turn lifecycle. Don't call `response.create`, and don't wait for the client to commit a user turn before translation begins. For browser media, use WebRTC. For server media pipelines such as phone calls or broadcast ingest, use WebSockets.
140
141See [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation) for the dedicated endpoint, session configuration, and architecture patterns.
142
143### Transcription sessions
144
145You can transcribe audio in more than one way. Use a realtime transcription session when your application needs live transcript deltas from streaming audio. Use the [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text) guide for file uploads, request-based transcription, translation, or speaker-labeling workflows.
146
147For realtime transcription, [`gpt-live-transcribe`](https://developers.openai.com/api/docs/models/gpt-live-transcribe) gives you controllable latency. Lower delay settings produce earlier partial text, while higher delay settings can improve transcript quality. Test with your real audio conditions, target languages, accents, and domain vocabulary before choosing a production default.
148 6
149See [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription) for session configuration and event handling.7For full-duplex conversations with a separate delegated backend, see [GPT-Live](https://developers.openai.com/api/docs/guides/live). To compare voice architectures and chained pipelines, see [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents).
150 8
151## Choose a connection method9## Build a speech-to-speech voice agent
152 10
153Choose the transport based on where your application captures and plays audio:11Use the Realtime API when the interaction should feel conversational and immediate. This is the best starting point for voice agents that need barge-in, low first-audio latency, natural turn taking, and realtime tool use.
154 12
155[WebRTC13The usual browser flow is:
156 14
151. Your application server creates an ephemeral client secret for the Realtime session.
162. Your frontend creates a `RealtimeSession`.
173. The session connects over WebRTC in the browser or WebSocket on the server.
184. The agent handles audio turns, tools, interruptions, and handoffs inside that session.
157 19
20Start a realtime voice session
158 21
159 Use for browser and mobile clients that capture or play audio directly.](https://developers.openai.com/api/docs/guides/realtime-webrtc)22```javascript
23import { RealtimeAgent, RealtimeSession } from "@openai/agents/realtime";
160 24
161[WebSocket25const agent = new RealtimeAgent({
26 name: "Assistant",
27 instructions: "You are a helpful voice assistant.",
28});
162 29
30const session = new RealtimeSession(agent, {
31 model: "gpt-realtime-2.1",
32});
163 33
34await session.connect({
35 apiKey: "ek_...(ephemeral key from your server)",
36});
37```
164 38
165 Use when your server already receives raw audio from a media pipeline, call
166 system, or worker.](https://developers.openai.com/api/docs/guides/realtime-websocket)
167 39
168[SIP40From there, attach tools, handoffs, and guardrails to the `RealtimeAgent` the same way you would attach them to a text agent. Keep audio transport concerns in the session layer, and keep business logic in the agent definition.
169 41
42Start with the transport docs when you need lower-level control:
170 43
171 44- [Audio and voice overview](https://developers.openai.com/api/docs/guides/audio)
172 Use for telephony voice agents. Confirm model support before using SIP for45- [Realtime API with WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime)
173 translation or transcription.](https://developers.openai.com/api/docs/guides/realtime-sip)46- [Realtime API with WebSocket](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime)
174 47
175## Safety identifiers48## Safety identifiers
176 49
188- Use [`POST /v1/realtime/client_secrets`](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets/methods/create) to create ephemeral credentials for browser or mobile clients.61- Use [`POST /v1/realtime/client_secrets`](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets/methods/create) to create ephemeral credentials for browser or mobile clients.
189- Use `/v1/realtime/calls` when establishing WebRTC sessions.62- Use `/v1/realtime/calls` when establishing WebRTC sessions.
190- Update session and event shapes for the GA interface. In particular, set `session.type`, move output audio configuration under `session.audio.output`, and use the newer response event names like `response.output_text.delta`, `response.output_audio.delta`, and `response.output_audio_transcript.delta`.63- Update session and event shapes for the GA interface. In particular, set `session.type`, move output audio configuration under `session.audio.output`, and use the newer response event names like `response.output_text.delta`, `response.output_audio.delta`, and `response.output_audio_transcript.delta`.
191- If you are moving a speech-to-speech app forward, start from the [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) guide. If you are moving a transcription workflow forward, use [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription).64- If you are moving a speech-to-speech app forward, start from the [browser example](#build-a-speech-to-speech-voice-agent). If you are moving a transcription workflow forward, use [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription).
65
66See the [Realtime client events reference](https://developers.openai.com/api/reference/resources/realtime/client-events), [Realtime sessions reference](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets), and [browser example](#build-a-speech-to-speech-voice-agent) for the current GA flow.
67
68
69
70
71
72## Next steps
73
74- [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations): Configure sessions and handle audio, text, and events.
75- [Voice activity detection](https://developers.openai.com/api/docs/guides/realtime-vad): Configure automatic turn detection.
76- [Tools and MCP](https://developers.openai.com/api/docs/guides/realtime-mcp): Add functions, MCP servers, and connectors.
77- [Prompting voice models](https://developers.openai.com/api/docs/guides/voice-prompting): Use the guide for your Realtime model.
78- [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=realtime): Understand Realtime accounting and caching.
79- [Server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=realtime): Keep tool execution and session control on your server.
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
192 100
193See the [Realtime client events reference](https://developers.openai.com/api/reference/resources/realtime/client-events), [Realtime sessions reference](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets), and [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) guide for the current GA flow.
194 101
195## Related guides
196 102
197- [Realtime prompting guide](https://developers.openai.com/api/docs/guides/realtime-models-prompting): Prompt and tune Realtime voice models.103## Other audio workflows
198- [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations): Work with the Realtime session lifecycle.
199- [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation): Translate live speech with a dedicated translation session.
200- [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription): Stream live transcript deltas from audio.
201- [Realtime with tools](https://developers.openai.com/api/docs/guides/realtime-mcp): Connect function tools, MCP servers, and connectors to a Realtime session.
202- [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls): Control Realtime sessions from your server.
203- [Managing costs](https://developers.openai.com/api/docs/guides/realtime-costs): Track and optimize Realtime API usage.
204 104
205Use [Audio and speech](https://developers.openai.com/api/docs/guides/audio) for the core concepts behind105The workflow chooser and shared audio vocabulary now live in [Audio and voice](https://developers.openai.com/api/docs/guides/audio). For continuous translation, use [Live translation](https://developers.openai.com/api/docs/guides/realtime-translation). For live captions, use [Live transcription](https://developers.openai.com/api/docs/guides/realtime-transcription); for recorded audio, use [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text).
206 audio input, audio output, streaming, latency, transcripts, and speech
207 generation. Use this overview when you are ready to choose an implementation
208 path.