1# Realtime transcription1# Realtime transcription
2 2
3You can use the Realtime API for transcription-only use cases, either with input from a microphone or from a file. For example, you can use it to generate subtitles or transcripts in real-time.3import {
4With the transcription-only mode, the model will not generate responses.4 Bolt,
5 5 Cube,
6If you want the model to produce responses, you can use the Realtime API in6 Desktop,
7 [speech-to-speech conversation mode](https://developers.openai.com/api/docs/guides/realtime-conversations).7 Phone,
8 8} from "@components/react/oai/platform/ui/Icon.react";
9## Realtime transcription sessions9
10 10Use realtime transcription when your application needs live speech-to-text without a spoken assistant response. Realtime transcription sessions stream transcript deltas as audio arrives, so users can see text before the full utterance is complete.
11To use the Realtime API for transcription, you need to create a transcription session, connecting via [WebSockets](https://developers.openai.com/api/docs/guides/realtime?use-case=transcription#connect-with-websockets) or [WebRTC](https://developers.openai.com/api/docs/guides/realtime?use-case=transcription#connect-with-webrtc).11
12 12For the lowest-latency streaming transcription path, use [`gpt-realtime-whisper`](https://developers.openai.com/api/docs/models/gpt-realtime-whisper). For offline files or workflows that don't need streaming deltas, use the standard speech-to-text models in the Audio API.
13Unlike the regular Realtime API sessions for conversations, the transcription sessions typically don't contain responses from the model.13
14 14## Choose a transcription model
15The transcription session object uses the same base session shape, but it always has a `type` of `"transcription"`:15
16<table>
17 <thead>
18 <tr>
19 <th>Model</th>
20 <th>Best for</th>
21 <th>Notes</th>
22 </tr>
23 </thead>
24 <tbody>
25 <tr>
26 <td className="whitespace-nowrap">
27 <a href="/api/docs/models/gpt-realtime-whisper">
28 gpt-realtime-whisper
29 </a>
30 </td>
31 <td>Live audio, transcript deltas, tunable latency.</td>
32 <td>Natively streaming and designed for realtime sessions.</td>
33 </tr>
34 <tr>
35 <td className="whitespace-nowrap">
36 <a href="/api/docs/models/gpt-4o-transcribe">gpt-4o-transcribe</a>
37 </td>
38 <td>Higher-accuracy speech-to-text where streaming isn't required.</td>
39 <td>Use for file and request-response transcription workflows.</td>
40 </tr>
41 <tr>
42 <td className="whitespace-nowrap">
43 <a href="/api/docs/models/gpt-4o-mini-transcribe">
44 gpt-4o-mini-transcribe
45 </a>
46 </td>
47 <td>Lower-cost transcription.</td>
48 <td>Use when cost matters more than top accuracy.</td>
49 </tr>
50 <tr>
51 <td className="whitespace-nowrap">
52 <a href="/api/docs/models/whisper-1">whisper-1</a>
53 </td>
54 <td>Existing Whisper integrations.</td>
55 <td>
56 Not natively streaming in the same way as{" "}
57 <code>gpt-realtime-whisper</code>.
58 </td>
59 </tr>
60 </tbody>
61</table>
62
63`gpt-realtime-whisper` is an alternative for live transcription, not a blanket replacement for every transcription model. Test it against your audio, languages, vocabulary, and latency requirements before switching production traffic.
64
65## Create a transcription session
66
67Realtime transcription uses a session with `type: "transcription"`. You can connect with [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket) for server-side audio pipelines or [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc) for browser audio.
16 68
17```json69```json
18{70{
19 "object": "realtime.session",71 "type": "session.update",
72 "session": {
20 "type": "transcription",73 "type": "transcription",
21 "id": "session_abc123",
22 "audio": {74 "audio": {
23 "input": {75 "input": {
24 "format": {76 "format": {
25 "type": "audio/pcm",77 "type": "audio/pcm",
26 "rate": 2400078 "rate": 24000
27 },79 },
28 "noise_reduction": {
29 "type": "near_field"
30 },
31 "transcription": {80 "transcription": {
32 "model": "gpt-4o-transcribe",81 "model": "gpt-realtime-whisper",
33 "prompt": "",
34 "language": "en"82 "language": "en"
35 },83 },
36 "turn_detection": {84 "turn_detection": {
40 "silence_duration_ms": 50088 "silence_duration_ms": 500
41 }89 }
42 }90 }
43 },91 }
44 "include": ["item.input_audio_transcription.logprobs"]92 }
45}93}
46```94```
47 95
48### Session fields96### Session fields
49 97
50- `type`: Always `transcription` for realtime transcription sessions.98- `type`: Set to `transcription` for transcription-only sessions.
51- `audio.input.format`: Input encoding for audio that you append to the buffer. Supported types are:99- `audio.input.format`: Input encoding for audio appended to the buffer. Use 24 kHz mono PCM when sending `audio/pcm`.
52 - `audio/pcm` (24 kHz mono PCM; only a `rate` of `24000` is supported).100- `audio.input.transcription.model`: Use `gpt-realtime-whisper` for streaming transcription.
53 - `audio/pcmu` (G.711 μ-law).101- `audio.input.transcription.language`: Optional language hint such as `en`.
54 - `audio/pcma` (G.711 A-law).102- `audio.input.turn_detection`: Optional voice activity detection. Set it to `null` if you want to commit audio manually.
55- `audio.input.noise_reduction`: Optional noise reduction that runs before VAD and turn detection. Use `{ "type": "near_field" }`, `{ "type": "far_field" }`, or `null` to disable.103
56- `audio.input.transcription`: Optional asynchronous transcription of input audio. Supply:104## Stream audio
57 - `model`: One of `whisper-1`, `gpt-4o-transcribe-latest`, `gpt-4o-mini-transcribe`, or `gpt-4o-transcribe`.105
58 - `language`: ISO-639-1 code such as `en`.106Send audio chunks with `input_audio_buffer.append`:
59 - `prompt`: Prompt text or keyword list (model-dependent) that guides the transcription output.107
60- `audio.input.turn_detection`: Optional automatic voice activity detection (VAD). Set to `null` to manage turn boundaries manually. For `server_vad`, you can tune `threshold`, `prefix_padding_ms`, `silence_duration_ms`, `interrupt_response`, `create_response`, and `idle_timeout_ms`. For `semantic_vad`, configure `eagerness`, `interrupt_response`, and `create_response`.108```javascript
61- `include`: Optional list of additional fields to stream back on events (for example `item.input_audio_transcription.logprobs`).109ws.send(
110 JSON.stringify({
111 type: "input_audio_buffer.append",
112 audio: base64Pcm16,
113 })
114);
115```
116
117If you disable turn detection, commit the buffer when you want transcription to begin:
118
119```javascript
120ws.send(
121 JSON.stringify({
122 type: "input_audio_buffer.commit",
123 })
124);
125```
126
127With server VAD enabled, the session commits audio automatically when it detects a turn boundary.
62 128
63You can find more information about the transcription session object in the [API reference](https://developers.openai.com/api/docs/api-reference/realtime-sessions/transcription_session_object).129## Handle transcript events
64 130
65## Handling transcriptions131Listen for incremental transcript deltas and completion events:
66 132
67When using the Realtime API for transcription, you can listen for the `conversation.item.input_audio_transcription.delta` and `conversation.item.input_audio_transcription.completed` events.133```javascript
134ws.on("message", (data) => {
135 const event = JSON.parse(data);
68 136
69For `whisper-1` the `delta` event will contain full turn transcript, same as `completed` event. For `gpt-4o-transcribe` and `gpt-4o-mini-transcribe` the `delta` event will contain incremental transcripts as they are streamed out from the model.137 if (event.type === "conversation.item.input_audio_transcription.delta") {
138 process.stdout.write(event.delta);
139 }
70 140
71Here is an example transcription delta event:141 if (event.type === "conversation.item.input_audio_transcription.completed") {
142 console.log("\nFinal transcript:", event.transcript);
143 }
144});
145```
146
147A delta event contains newly available transcript text:
72 148
73```json149```json
74{150{
75 "event_id": "event_2122",
76 "type": "conversation.item.input_audio_transcription.delta",151 "type": "conversation.item.input_audio_transcription.delta",
77 "item_id": "item_003",152 "item_id": "item_003",
78 "content_index": 0,153 "content_index": 0,
80}155}
81```156```
82 157
83Here is an example transcription completion event:158A completion event contains the final transcript for the committed item:
84 159
85```json160```json
86{161{
87 "event_id": "event_2122",
88 "type": "conversation.item.input_audio_transcription.completed",162 "type": "conversation.item.input_audio_transcription.completed",
89 "item_id": "item_003",163 "item_id": "item_003",
90 "content_index": 0,164 "content_index": 0,
92}166}
93```167```
94 168
95Note that ordering between completion events from different speech turns is not guaranteed. You should use `item_id` to match these events to the `input_audio_buffer.committed` events and use `input_audio_buffer.committed.previous_item_id` to handle the ordering.169Ordering between completion events from different speech turns isn't guaranteed. Use `item_id` to match transcription events to committed input items.
96
97To send audio data to the transcription session, you can use the `input_audio_buffer.append` event.
98
99You have 2 options:
100
101- Use a streaming microphone input
102- Stream data from a wav file
103
104{/*
105
106### Using microphone input
107
108
109
110<div data-content-switcher-pane data-value="js">
111 <div class="hidden">ws module (Node.js)</div>
112 </div>
113 <div data-content-switcher-pane data-value="python" hidden>
114 <div class="hidden">websocket-client (Python)</div>
115 </div>
116
117
118
119### Using file input
120
121 170
171## Tune latency and accuracy
122 172
123<div data-content-switcher-pane data-value="js">173Streaming transcription trades latency for transcript quality. Lower delay settings can produce earlier partial text. Higher delay settings give the model more audio context before emitting text and can improve word error rate.
124 <div class="hidden">ws module (Node.js)</div>
125 </div>
126 <div data-content-switcher-pane data-value="python" hidden>
127 <div class="hidden">websocket-client (Python)</div>
128 </div>
129 174
175Start by testing a few delay targets against your real audio. Useful evaluation points are:
130 176
131*/}177- 0.4 seconds for the most latency-sensitive interactions;
132## Voice activity detection178- 0.8 to 1.2 seconds for balanced live captions;
179- 1.5 to 2.0 seconds when accuracy matters more than immediate display;
180- 3.0 seconds for workflows that can tolerate more delay.
133 181
134The Realtime API supports automatic voice activity detection (VAD). Enabled by default, VAD will control when the input audio buffer is committed, therefore when transcription begins.182Don't choose a setting from synthetic audio alone. Test with representative microphones, telephony audio, accents, background noise, code-switching, domain vocabulary, and long sessions.
135 183
136Read more about configuring VAD in our [Voice Activity Detection](https://developers.openai.com/api/docs/guides/realtime-vad) guide.184## Guide vocabulary and domain terms
137 185
138You can also disable VAD by setting the `audio.input.turn_detection` property to `null`, and control when to commit the input audio on your end.186If your application depends on exact domain vocabulary, include a language hint and test whether your model and endpoint support prompt or keyword steering before relying on it. Where supported, use short keyword lists rather than long instructions.
139 187
140## Additional configurations188Example keyword style:
141 189
142### Noise reduction190```text
143 191Keywords: metoprolol, atorvastatin, A1C, systolic, diastolic
144Use the `audio.input.noise_reduction` property to configure how to handle noise reduction in the audio stream.192```
145 193
146- `{ "type": "near_field" }`: Use near-field noise reduction (default).194For production, treat keyword steering as an aid rather than a guarantee. Continue to evaluate names, numbers, dates, medication names, product names, artist names, and other high-value entities manually.
147- `{ "type": "far_field" }`: Use far-field noise reduction.
148- `null`: Disable noise reduction.
149 195
150### Using logprobs196## Handle confidence, timestamps, and diarization
151 197
152You can use the `include` property to include logprobs in the transcription events, using `item.input_audio_transcription.logprobs`.198Only request optional fields that your selected model and endpoint support. If your application needs confidence scoring, timestamps, or diarization, verify support before launch and add fallbacks for fields that aren't available.
153 199
154Those logprobs can be used to calculate the confidence score of the transcription.200When log probabilities are available, request them with `include`:
155 201
156```json202```json
157{203{
158 "type": "session.update",204 "type": "session.update",
159 "session": {205 "session": {
206 "type": "transcription",
160 "audio": {207 "audio": {
161 "input": {208 "input": {
162 "format": {
163 "type": "audio/pcm",
164 "rate": 24000
165 },
166 "transcription": {209 "transcription": {
167 "model": "gpt-4o-transcribe"210 "model": "gpt-realtime-whisper"
168 },
169 "turn_detection": {
170 "type": "server_vad",
171 "threshold": 0.5,
172 "prefix_padding_ms": 300,
173 "silence_duration_ms": 500
174 }211 }
175 }212 }
176 },213 },
178 }215 }
179}216}
180```217```
218
219## Production checklist
220
221- Pick a target latency and accuracy threshold before tuning.
222- Test against real production audio, not only clean samples.
223- Test each target language.
224- Include numbers, dates, currency, email addresses, product names, and domain terms in your eval set.
225- Track empty, truncated, and delayed transcripts apart from word error rate.
226- Decide how your UI should revise partial text when later deltas correct earlier text.
227- Use `item_id` to order and reconcile final transcripts.
228- Keep a fallback path for unsupported timestamps, diarization, or confidence fields.
229
230## Related guides
231
232<a href="/api/docs/guides/realtime">
233
234
235<span slot="icon">
236 </span>
237 Compare voice-agent, translation, and transcription sessions.
238
239
240</a>
241
242<a href="/api/docs/guides/realtime-translation">
243
244
245<span slot="icon">
246 </span>
247 Translate live speech with a dedicated translation session.
248
249
250</a>
251
252<a href="/api/docs/guides/realtime-websocket">
253
254
255<span slot="icon">
256 </span>
257 Stream raw audio through a server-side media pipeline.
258
259
260</a>
261
262<a href="/api/docs/guides/realtime-vad">
263
264
265<span slot="icon">
266 </span>
267 Configure turn detection for live audio streams.
268
269
270</a>