2 2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4 4
5After [connecting to GPT-Live](https://developers.openai.com/api/docs/guides/live), use session events to update context, display transcripts, and manage the connection's lifecycle. The model can listen and speak at the same time, so keep received events, audio playback, and backend task state separate in your application.5After [connecting to GPT-Live](https://developers.openai.com/api/docs/guides/live), use session events to add context, display transcripts, and manage the connection. GPT-Live can listen and speak at the same time. Track transcript text, played audio, and backend task progress separately so your interface can show what the assistant is saying and what work is still running.
6 6
7This guide assumes your connection has emitted `session.started`. See [Connections](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for connection setup and audio streaming, and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for backend work.7This guide assumes your connection has emitted `session.started`. See [Connections](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for connection setup and audio streaming, and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for backend work.
8 8
20| ------------ | --------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |20| ------------ | --------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |
21| Model | Set the required `model`. | Start a new session to change it. |21| Model | Set the required `model`. | Start a new session to change it. |
22| Instructions | Set `instructions` for conversation behavior, up to 16,384 tokens. | Add instructions with `session.instructions.append`. |22| Instructions | Set `instructions` for conversation behavior, up to 16,384 tokens. | Add instructions with `session.instructions.append`. |
23| History | Set `input` to relevant prior text messages. It defaults to `[]`. | Append context; don't replace the startup history. |23| History | Set `input` to relevant prior text messages. It defaults to `[]`. | Add context with append events. |
24| Voice | Set `audio.output.voice` to a supported voice or authorized custom voice. The default is `marin`. | Start a new session to change it. |24| Voice | Set `audio.output.voice` to a supported voice or authorized custom voice. The default is `marin`. | Start a new session to change it. |
25| Delegation | Set `delegation.type` to `client` or `responses`. Omitted or `null` delegation selects client mode. | Update Responses settings within the existing mode. |25| Delegation | Set `delegation.type` to `client` or `responses`. Omitted or `null` delegation selects client mode. | Update Responses settings within the existing mode. |
26| Storage | Set `store` to `true` to make the session available for forking. It defaults to `false`. | Choose at startup. |26| Storage | Set `store` to `true` to make the session available for forking. It defaults to `false`. | Choose at startup. |
44| Delta | `delta` | English | Southern U.S. | Feminine | Generated |44| Delta | `delta` | English | Southern U.S. | Feminine | Generated |
45| Cinder | `cinder` | English | Southern U.S. | Masculine | Generated |45| Cinder | `cinder` | English | Southern U.S. | Masculine | Generated |
46 46
47Regional influence describes a voice's speaking style, not a guarantee of accent fidelity. For an approved voice created from your own recording, see [Custom voices](https://developers.openai.com/api/docs/guides/custom-voices).47Regional influence describes a voice’s speaking style. Test the voice with the languages and pronunciation your application needs. For an approved voice created from your own recording, see [Custom voices](https://developers.openai.com/api/docs/guides/custom-voices).
48 48
49 49
50 50
51 51
52 52
53For WebSocket, choose the shared `audio.format` at startup; it cannot change during the session. For WebRTC, omit this field because the connection negotiates its audio format. See [WebSocket audio formats](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for format and streaming details.53For WebSocket, choose `audio.format` at startup. The same format applies to input and output audio for the session. To use another format, start a new session. WebRTC negotiates its audio format during connection setup, so leave `audio.format` out of WebRTC requests. See [WebSocket audio formats](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for supported formats and streaming details.
54 54
55### Update a live session55### Update a live session
56 56
57Use `session.update` for changes to `session.delegation.responses` in a session already using Responses delegation. Send only the settings you want to change; omitted settings retain their values. See [Configure Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for the settings and update workflow.57Use `session.update` for changes to `session.delegation.responses` in a session already using Responses delegation. Send only the settings you want to change; omitted settings retain their values. See [Configure Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for the settings and update workflow.
58 58
59You cannot change the delegation mode after startup. In particular, setting `delegation` to `null` selects client mode; it does not reset a Responses session. The startup fields `model`, `instructions`, `input`, `audio`, and `store` are not accepted update fields. Unknown configuration fields are rejected.59Choose the delegation mode and the fields `model`, `instructions`, `input`, `audio`, and `store` at startup. Use `session.update` only for the supported Responses settings described above; other configuration fields are rejected. To switch delegation modes, create a new session. At startup, `delegation: null` selects client delegation rather than restoring default Responses settings.
60 60
61A successful update emits `session.updated` with the full resolved session configuration. When you supply an `event_id`, the acknowledgment returns it as `client_event_id`. Check for [rejected commands](#handle-rejected-commands) as well as acknowledgments. Acceptance confirms the configuration update; it does not establish that a backend task ran or that the model spoke.61A successful update emits `session.updated` with the resulting session configuration. Match it to your outgoing `event_id` through `client_event_id`. Handle [rejected commands](#handle-rejected-commands) in the same event loop. Track backend work and spoken output through their own events.
62 62
63## Provide history and context63## Provide history and context
64 64
94```94```
95 95
96 96
97The list accepts up to 128 messages and 8,192 combined tokens. Supported roles are `developer`, `user`, and `assistant`, each with one text part. Developer and user messages use `input_text`; assistant messages use `text` or `output_text`. Put trusted application instructions in `instructions` or a developer message. The list does not accept the `system` role.97The list accepts up to 128 messages and 8,192 combined tokens. Each message has one text part and one of these roles: `developer`, `user`, or `assistant`. Developer and user messages use `input_text`; assistant messages use `text` or `output_text`. Put trusted application instructions in `instructions` or a developer message.
98 98
99Select the history needed for the next interaction. `input` is a startup field, not a way to replace history during a running session. It also does not accept the full range of backend input items used in Responses delegation.99Select the text history needed for the next interaction and supply it at startup. During the session, add updates with the context events below. Send backend-specific items, such as tool results, through the [delegation workflow](https://developers.openai.com/api/docs/guides/live-delegation).
100 100
101### Understand when context reaches the model101### Understand when context reaches the model
102 102
103The full `input` supplied at session creation is available to the model when the session starts. Put context the model needs from the beginning in this field.103Put any context the model needs from the start in `input`; the full field is available when the session starts.
104 104
105During a running session, the `session.instructions.append`, `session.thinking.append`, and `session.commentary.append` events feed content into the model over time. Their acknowledgments wait until frame progress reaches the estimated end of context injection. The returned `start_ms` and `end_ms` describe an estimated range on the session timeline, not speech or playback completion. They do not prove that the model consumed the entire update. Don't assume its next speech will reflect the whole update.105During a running session, `session.instructions.append`, `session.thinking.append`, and `session.commentary.append` add context over time. The acknowledgment arrives when the session timeline reaches the estimated end of the added context. Its `start_ms` and `end_ms` estimate where that update falls on the session timeline.
106 106
107If frame progress stops, an acknowledgment can remain pending. Closing the session reports an error for pending appends. Match each acknowledgment to the outgoing `event_id` through `client_event_id`, and keep handling errors while you wait.107These times describe context delivery, not speech or playback. The model may still respond before it has used the whole update. When an action depends on a new instruction or fact, verify the resulting behavior in your application.
108
109If the session timeline stops, the acknowledgment can remain pending. Match acknowledgments to the outgoing `event_id` through `client_event_id`, and keep handling errors while you wait. Closing the session returns errors for appends that are still pending.
108 110
109### Add context during the conversation111### Add context during the conversation
110 112
147```149```
148 150
149 151
150Wait for `session.thinking.appended` with `client_event_id: "context_1"`, or handle an error. The acknowledgment confirms that context was accepted. It does not confirm speech, playback, or completion of an external action.152Handle `session.thinking.appended` with `client_event_id: "context_1"`, or the corresponding error, to track this update. See [Understand when context reaches the model](#understand-when-context-reaches-the-model) for acknowledgment timing.
151 153
152Quiet context can influence later speech; it is not a privacy boundary. Keep credentials, secrets, and text the model must never reveal out of all three events. Use the instructions event for application-authored behavior, not untrusted tool output. Enforce permissions and required confirmations in your application.154The assistant may repeat information supplied through any of these events. Send only information suitable for the conversation, and keep credentials and secrets in your backend. Use `session.instructions.append` for behavior defined by your application. Supply factual tool results as context, and enforce permissions and required confirmations in application code.
153 155
154For page navigation, selections, and other UI changes, see [Share UI context](https://developers.openai.com/api/docs/guides/live-delegation#share-ui-context) for concise updates that help GPT-Live understand what the user is referring to.156For page navigation, selections, and other UI changes, see [Share UI context](https://developers.openai.com/api/docs/guides/live-delegation#share-ui-context) for concise updates that help GPT-Live understand what the user is referring to.
155 157
156For results tied to a backend task, use a known client delegation ID and follow [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update). That ID is not a Responses response ID or tool call ID.158For an update about a specific backend task, use the ID of the relevant client delegation. A delegation ID identifies the Live task; Responses response IDs and tool call IDs identify different objects. See [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update) for the workflow.
159
160When your application detects a problem, send a short correction through the session’s primary WebSocket or a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband). See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for checks, action controls, and playback handling.
161
162
157 163
158Use instructions to steer the conversation after an application check triggers. Your server can monitor events and send these corrections through a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) attached to the existing session, or through its primary WebSocket. See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for concurrent checks, action blocking, and playback control.
159 164
160 165
161 166
164 169
165 170
166 171
167## Build the conversation interface
168 172
169Display transcripts and microphone state independently of backend progress. Receiving assistant text does not tell you how much audio the user has heard.173## Manage speech and transcripts
170 174
171### Transcript deltas175### Transcript deltas
172 176
182}186}
183```187```
184 188
185Append fragments in order for each speaker, retaining `start_ms` and `end_ms`. These are milliseconds on the session timeline, with intervals that include the start and exclude the end. They are not wall-clock timestamps, packet arrival times, or exact word alignments.189Append each speaker’s `delta` fragments exactly as received, preserving spaces and repeated words. Retain their `start_ms` and `end_ms`. These values are milliseconds from the start of the session. The example above covers the interval from 1,000 ms up to, but excluding, 1,200 ms. They describe approximate fragment timing rather than exact word boundaries; use them instead of packet arrival times to organize the transcript.
186 190
187Only intervals containing transcript text produce events, and network delivery can be uneven. Do not infer silence from a missing event or treat a fragment as a complete user turn. Transcript deltas have no item ID or authoritative turn-completed event.191Transcript events arrive for intervals that contain text, and delivery can be uneven. A fragment may contain only part of a sentence; a gap in delivery may be a network delay. Transcript deltas have no item ID or event that marks a completed conversational turn, so your application decides how to group them for display.
188 192
189Processing transcript fragments is optional. You can use them to update your UI, run checks, or start work early while the conversation continues. For lightweight checks, consider a small model such as `gpt-5.6-luna` with low reasoning effort. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples and connection guidance.193Processing transcript fragments is optional. You can use them to update your UI, run checks, or start work early while the conversation continues. For lightweight checks, consider a small model such as `gpt-5.6-luna` with low reasoning effort. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples and connection guidance.
190 194
191For conversation guardrails, check accumulated user and assistant text as it arrives. Transcript delivery does not provide an advance buffer for approving speech before playback. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).195Use [transcript guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#run-checks-alongside-the-conversation) to monitor the conversation and trigger interventions while speech continues. If your application needs to check assistant speech before playback, see [Check speech before playback](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#check-speech-before-playback) for buffering, approval, interruption, and recovery handling.
192 196
193 197
194 198
195 199
196 200
197If your interface groups text into turns, keep that grouping revisable. Preserve the original fragments, allow user and assistant intervals to overlap, and tune any gap timeout against recorded conversations. A brief acknowledgment from the other speaker may belong within an ongoing exchange. Grouping fragments must not trigger tool execution or cancel backend work by itself.201Keep transcript timing separate from audio playback. WebSocket `session.output_audio.delta` events have no timing fields or output-audio-done event; WebRTC delivers audio through its media track. See [Connections](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for audio handling.
198
199 202
200 203
201 204
202 205
203Keep transcript timing separate from audio playback. WebSocket `session.output_audio.delta` events have no timing fields or output-audio-done event; WebRTC delivers audio through its media track. See [Connections](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for audio handling.
204 206
205### Display captions207### Display captions
206 208
207Build caption rows that can grow while both speakers are talking:209GPT-Live is full duplex: the caller and assistant can speak at the same time. Update their captions independently so both speakers’ text can keep growing during overlapping speech.
208 210
2091. **Preserve the text.** Store each speaker's original `delta`, `start_ms`, and `end_ms`. Concatenate text exactly as received, including spaces and repeated words. Don't trim fragments or insert spaces between them.211If your app uses chat bubbles, the fragments “I’d like” and “ to change my booking” can appear in one caller bubble. If the assistant says “Sure” while the caller continues, show that acknowledgment separately while allowing the caller’s bubble to keep growing. Keep the original fragments and timestamps so text that arrives later can update the appropriate bubble.
2102. **Update each speaker independently.** Allow user and assistant rows to grow during overlapping speech. Keep earlier assistant text visible after an interruption, and start a new row when the assistant resumes.
2113. **Keep rows stable.** Assign display IDs in your application and preserve row order as text grows. Don't derive row identity from changing text or end timestamps, or move a row to the bottom whenever it receives a fragment.
2124. **Revisit grouping for late fragments.** Use transcript timestamps to group nearby fragments from the same speaker. Allow late text to update earlier rows and revise fragment assignments while retaining the original fragments. These display groups are not complete semantic turns; any gap threshold is an application choice to test.
2135. **Let the reader control scrolling.** Follow new text while the reader is at the bottom. Pause automatic scrolling when they scroll up, and provide a way to return to the latest captions.
2146. **Show tool progress in a status area.** Use assistant transcript events for spoken captions. Display tool activity and backend results outside the captions; receiving a result does not mean the assistant has said it.
215 212
216Test the display with overlapping speech, short acknowledgments, interruptions, long pauses, and translation where the two speakers' text arrives at different rates.213Use `session.output_transcript.delta` for spoken captions and show backend updates separately. Keep decisions about running tools or canceling work in your application’s task logic, separate from how you group text for display.
217 214
218### Control microphone input215### Control microphone input
219 216
244 241
245Wait for `session.input_audio.muted` with `client_event_id: "mute_1"` before treating the command as accepted. To resume input, send `session.input_audio.unmute` and wait for `session.input_audio.unmuted`. Handle errors for either command.242Wait for `session.input_audio.muted` with `client_event_id: "mute_1"` before treating the command as accepted. To resume input, send `session.input_audio.unmute` and wait for `session.input_audio.unmuted`. Handle errors for either command.
246 243
247Muting input does not stop inference, delegated work, or generated speech. Control microphone capture and audio playback separately in your application when those controls are needed.244Muting input leaves the session running: the model can keep generating speech, and delegated work can continue. Use your application’s microphone capture and audio player controls when you also need to stop local recording or playback.
248 245
249### Greet before the caller speaks246### Greet before the caller speaks
250 247
251To request a greeting after `session.started`:248To have GPT-Live open the conversation, send greeting instructions after `session.started`. Specify the language, what the assistant should say, and that it should begin immediately, then pause to listen. Use the application’s chosen greeting language until the caller speaks. For example:
252 249
2531. Send one fresh `session.instructions.append` with `delegation_id: null`. Include the greeting, its language, and an explicit instruction to greet immediately without waiting for the caller, then pause and listen. Keep the existing startup instructions.250> Greet the caller now in English. Introduce yourself as the support assistant and ask how you can help. Then pause and listen.
2542. Wait for `session.instructions.appended`, matching its `client_event_id` to your command. Handle a rejected command before continuing.
2553. Keep input audio running, including silence before the caller speaks. On WebSocket, continue sending `session.input_audio.append`; on WebRTC, keep the negotiated input audio track active. Observe output transcript and audio for the greeting.
256 251
257Use the language specified by your application until the caller speaks; don't infer it from a name, phone number, or location. See [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for prompt design.2521. Keep input audio running throughout this sequence, including silence before the caller speaks. On WebSocket, continue sending `session.input_audio.append`; on WebRTC, keep the input audio track active.
2532. Send the instructions once with `session.instructions.append` and `delegation_id: null`.
2543. Match `session.instructions.appended` to your command using `client_event_id`, and handle any error. This acknowledgment confirms that the instructions were accepted.
258 255
259For a greeting that needs to follow application instructions, send those instructions with `session.instructions.append`, then use a short `session.commentary.append` to prompt the assistant to begin. For example: “Begin the conversation now, following the instructions provided.” Keep input audio running, including silence before the caller speaks.256For exact wording and a known playback-completion point, play a verified recording or rendered clip through your application and [control GPT-Live playback while it plays](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed). Test greetings in the languages you support, including when the caller starts speaking during the greeting. See [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for prompt design.
260
261Instructions request a greeting; they do not guarantee exact wording or uninterrupted playback. The API does not emit an opening-completed event, and acknowledgment does not mean the greeting was heard. Use application-controlled playback if the audio must be verbatim. Test your greeting with the languages and interruptions your application supports.
262 257
263### Deliver a disclosure258### Deliver a disclosure
264 259
296```291```
297 292
298 293
299Keep input audio running as described in [Greet before the caller speaks](#greet-before-the-caller-speaks). Choose the delivery point deliberately: an instruction sent during the conversation can interrupt speech in progress.294Keep input audio running, as in [Greet before the caller speaks](#greet-before-the-caller-speaks). An instruction sent during the conversation can interrupt speech in progress.
300 295
301This requests the wording; it does not guarantee exact delivery. Verify the complete spoken disclosure and actual playback before marking it delivered. `session.instructions.appended` confirms only that the instruction was accepted. If exact audio delivery is required, play a verified recording or rendered clip through your application and control GPT-Live output while it plays. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).296Check the generated disclosure and its playback before marking it delivered. The instruction acknowledgment records acceptance; use the audio itself to check the wording. For exact wording and a known playback-completion point, play a verified recording or rendered clip through your application and control GPT-Live output while it plays. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).
302 297
303 298
304 299
306 301
307## Manage longer conversations302## Manage longer conversations
308 303
309GPT-Live automatically manages context during long conversations; no configuration parameter is needed. The instructions you provide at session start are preserved throughout compaction. You don’t need to resend them.304GPT-Live manages long conversations automatically and preserves your original startup instructions.
310 305
311The default context window holds 128,000 tokens, including your instructions, conversation text, and audio tokens that don’t appear in the transcript.306The default context window holds 128,000 tokens, including your instructions, conversation text, and audio tokens that don’t appear in the transcript.
312 307
320 315
321## Store and fork a session316## Store and fork a session
322 317
323Set `store` to `true` in the session configuration at creation to save a recording for later download or forking. Storage defaults to `false` and must be enabled for your project. Downloads and forks require a completed stored recording and a data policy that permits persistence. Recordings expire after 30 days. With Zero Data Retention, `store` is treated as `false`, and recording downloads and forks are unavailable. See [GPT-Live data controls](https://developers.openai.com/api/docs/guides/your-data#v1livesessions).318A fork starts a new session from a saved voice conversation. Use it to run several evaluation trials from the same reference conversation, or to let a user continue after an earlier session has ended. Each fork gets a new connection and session ID, with the source conversation and its saved configuration as its starting point.
319
320### Run evaluations from a reference conversation
321
322Suppose you want to test how your agent handles a caller changing an order. Record the setup once, through the point where the caller has identified the order. End and finalize that session before the caller asks to change it. Each evaluation can then fork the same reference session and receive the same next caller audio: “Actually, can you send it to my office instead?”
323
324For each trial, restore the same test order and application state, supply the next caller input, and evaluate the new response and tool actions. You can repeat the scenario or compare supported Responses backend settings. Measure fork startup separately from response time. See the [GPT-Live evaluation guide](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) for choosing scenarios and measuring results.
325
326Forks inherit the GPT-Live model, voice, and original instructions. To compare a different voice model or startup prompt, create new sessions with that configuration. The fork API uses the completed source recording, so end the reference session where you want the evaluation to begin.
327
328### Continue after a session ends
324 329
325For example, add this field to the `session` object in your WebSocket `session.start` event or WebRTC creation request:330For example, a caller may hang up and call back later, or reconnect after a dropped call. If the earlier session has a completed stored recording, your application can fork it on a new connection and continue from the saved conversation.
331
332Save application task state alongside the source session ID. Before continuing, check the status of any outstanding backend work and give the new session its current results. For example, if an order update was already submitted, confirm its outcome before attempting another update. Use the new session ID for controls and sideband connections, and route subsequent backend results to the new session.
333
334### Prepare a session for forking
335
3361. **Enable storage when you create the source.** Set `store: true` in its session configuration. Storage defaults to `false`, must be enabled for your project, and requires a data policy that permits persistence.
3372. **Save the source session ID.** Read it from `session.started` or the WebRTC creation response, and associate it with your application’s conversation record.
3383. **Finish and close the source.** Complete required backend work, then follow [Usage and graceful close](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close). Keep the connection open until `session.closed` and handle any finalization error. Forking requires a completed stored recording; saving it can add time to finalization.
3394. **Start a fork on a new connection.** Use the source ID with the transport flow below, save the new session ID, and complete startup before continuing the conversation. Set the child’s `store` explicitly: `true` if you want to fork its continuation later, or `false` if you do not need to store that trial. Omitting it inherits the source setting.
340
341Stored recordings are available for 30 days. With Zero Data Retention (ZDR), `store` is treated as `false` and forking is unavailable. For fork-based evaluations, use a non-ZDR organization with storage enabled for the project.
342
343If you have no completed stored recording, [start a new session with relevant saved text history](https://developers.openai.com/api/docs/guides/live-conversations#seed-a-session-with-prior-conversation). See [GPT-Live data controls](https://developers.openai.com/api/docs/guides/your-data#v1livesessions) for storage requirements.
344
345For example, set this field in the source session’s WebSocket `session.start` configuration or WebRTC creation request:
326 346
327```json347```json
328{348{
330}350}
331```351```
332 352
333Save the source session ID from `session.started` or the WebRTC creation response. A fork starts a **new session with a new ID** from the stored session state. It does not reopen the original connection or reuse the source session ID.
334
335Start the fork through the transport your application uses:353Start the fork through the transport your application uses:
336 354
337| Transport | Start the fork |355| Transport | Start the fork |
339| WebSocket | Connect to `wss://api.openai.com/v1/live/sessions/{source_session_id}/fork`. |357| WebSocket | Connect to `wss://api.openai.com/v1/live/sessions/{source_session_id}/fork`. |
340| WebRTC | Send a new SDP offer to `POST /v1/live/sessions/{source_session_id}/fork`. Apply the returned `transport.sdp` answer to the new peer connection. |358| WebRTC | Send a new SDP offer to `POST /v1/live/sessions/{source_session_id}/fork`. Apply the returned `transport.sdp` answer to the new peer connection. |
341 359
342A fork inherits the stored session configuration, subject to the transport rules below. For a WebSocket fork, send `session.start` with a required `session` object; `{}` supplies no overrides. Do not supply a new model or repeat the original instructions or input. You can override `store`, Responses delegation settings, and the new WebSocket audio format. WebRTC forks can override `store`, Responses delegation settings, and frontend client permissions. Omitting `store` on a fork inherits the source session's setting.360A fork inherits the source session’s model, original instructions, and input. Send only the supported overrides at startup:
361
362- **WebSocket:** `store`, Responses delegation settings, and the new connection’s `audio.format`. Send a `session.start` event with a `session` object; use `{}` to keep inherited settings where supported.
363- **WebRTC:** `store`, Responses delegation settings, and frontend client permissions.
343 364
344A WebSocket fork does **not** inherit the source audio format: set `audio.format` explicitly or use the default PCM16 at 24 kHz. It also discards inherited frontend data-channel permissions. WebRTC forks negotiate their audio format and reject `audio.format`; they preserve frontend permission settings unless you override them.365For a WebSocket fork, set `audio.format` for the new connection or use the default PCM16 at 24 kHz. The source audio format and frontend data-channel permissions are not inherited. WebRTC negotiates audio format during connection setup; omit `audio.format`. WebRTC preserves frontend permission settings unless you override them.
345 366
346Wait for `session.started` before sending further WebSocket commands. WebRTC starts through the HTTP request and must not receive a second `session.start` on its data channel.367For WebSocket, wait for `session.started` before sending more commands. For WebRTC, the HTTP request starts the session; continue through the negotiated connection without sending another `session.start`.
347 368
348### Start a WebSocket fork369### Start a WebSocket fork
349 370
438 459
439Return the response to your frontend, apply `transport.sdp` as the new peer connection's answer, and retain the new `session.id`. Keep the API key on your backend.460Return the response to your frontend, apply `transport.sdp` as the new peer connection's answer, and retain the new `session.id`. Keep the API key on your backend.
440 461
441Use the new session ID for later sideband connections and session controls. Keep application task state separately: restoring conversation state does not confirm that a pending backend action completed. Reconcile uncertain results before retrying an action. If you don't have a stored session to fork, [seed a new session with saved history](#seed-a-session-with-prior-conversation).462Use the new session ID for sideband connections and session controls. Before retrying an unfinished action, check its outcome in your backend and restore the current application task state. If you have no completed stored recording, [seed a new session with saved history](#seed-a-session-with-prior-conversation).
442 463
443### Download a recording464### Download a recording
444 465
470```491```
471 492
472 493
494## Close idle sessions and resume
495
496For applications with long gaps between interactions, close the voice session during inactivity and start a new session when the user returns. Keep conversation context and application task state so the user can continue without repeating themselves. For example, an in-car assistant can resume when the driver activates voice again, while a coding assistant can keep its backend worker running between voice conversations.
497
4981. **Decide when to close.** Use an application-controlled inactivity timeout based on audio activity, assistant playback, and application interactions. Allow for expected pauses, such as reading or thinking. Close only when playback has finished and no pending work requires the current voice session. Gaps between transcript events alone do not establish silence.
4992. **Save state and close gracefully.** Save the source session ID, conversation context, and current task state. Finish any required Responses work, then follow [Usage and graceful close](#usage-and-graceful-close): install the `session.closed` listener, send `session.close`, and wait for `session.closed` before releasing the connection. With client delegation, application-managed backend work can continue independently while voice is closed.
5003. **Detect when to restart.** Offer a button labeled **Resume conversation**, a push-to-talk control, or an application-managed wake trigger. A closed Live session cannot listen for the user. If you use local speech detection to restart automatically, keep microphone capture active and buffer the opening speech through connection setup. Deliver that audio once the new session is ready, so the user’s first words are preserved.
5014. **Restore context in a new session.** If the source was created with `store: true`, storage is enabled and permitted, and its recording finalized successfully, [fork the stored session](#prepare-a-session-for-forking). Otherwise, [start a new session with saved text history](#seed-a-session-with-prior-conversation). Save the new session ID, check the status of outstanding backend operations, and route subsequent results to the new session. Keep operation status in your application so restarting does not repeat completed actions.
502
503Muting the microphone leaves the session active. Choose an idle timeout by comparing avoided voice duration with session-creation costs and the delay before voice becomes ready again. See [Voice session costs](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live#voice-session-costs) and [WebRTC initialization charges](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live#webrtc-initialization-charges).
504
473## Handle errors and end the session505## Handle errors and end the session
474 506
475Keep reading session events until the session finalizes. Distinguish a rejected command, a failed connection, and a completed session so your application can recover appropriately.507Keep reading session events until the session finalizes. Distinguish a rejected command, a failed connection, and a completed session so your application can recover appropriately.
492}524}
493```525```
494 526
495An error code can be `null`, and an error may lack a client event ID. Handle those cases without assuming a command succeeded. For an immutable-field error, keep the current configuration or create a new session with the intended settings.527Provide a general error handler for errors whose code is `null` or whose client event ID is absent. For an immutable-field error, keep the current configuration or create a new session with the intended settings.
496 528
497### Handle moderation529### Handle moderation
498 530
501- Some moderation events end the session.533- Some moderation events end the session.
502- Others cut off assistant audio for the remainder of its current speech and emit an `error` event without ending the session.534- Others cut off assistant audio for the remainder of its current speech and emit an `error` event without ending the session.
503 535
504Read `error` events even while audio is playing. Don't assume every moderation error closes the session, or that an audio interruption means the connection failed. Keep application state aligned with the session lifecycle, and don't mark an interrupted spoken message as fully delivered. Application-level [conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) remain separate from this built-in moderation behavior.536Keep handling `error` events while audio is playing. Track audio interruption and session closure separately, and mark a spoken message as delivered only after checking its playback. Apply your own [conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) alongside built-in moderation.
505 537
506### Usage and graceful close538### Usage and graceful close
507 539
516}548}
517```549```
518 550
519These are snapshots, not increments to sum. Backend token usage is separate; preserve it from nested Responses completion events. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting.551Use the latest `usage.seconds` as the running total for voice duration. For example, updates of 12 and then 15 seconds mean 15 seconds of use. Track backend token usage separately from nested Responses completion events. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting.
520 552
521To close gracefully:553To close gracefully:
522 554
528 560
529Sending `session.close` cancels queued Responses and rejects further commands. An active response can finish, but one waiting for a function result cannot continue after closing starts. Decide separately whether to finish or cancel work your application runs through client delegation.561Sending `session.close` cancels queued Responses and rejects further commands. An active response can finish, but one waiting for a function result cannot continue after closing starts. Decide separately whether to finish or cancel work your application runs through client delegation.
530 562
531The `session.closed` event establishes finalization; the embedded session is a configuration snapshot. A socket close alone does not establish success, and a transport close code after a valid final event does not invalidate finalization. Closing WebRTC immediately after sending the command can prevent delivery of the final event.563Use `session.closed` to confirm finalization and read the final configuration snapshot. Keep the transport open until this event arrives. If the socket closes first, record finalization as unconfirmed; if it closes after a valid `session.closed`, retain the confirmed result.
532 564
533The final event's `reason` explains why the session ended:565The final event's `reason` explains why the session ended:
534 566
546 578
547An HTTP session-creation error means the session did not reach `session.started`. Handle startup errors separately from errors in a running session. If a running connection fails before `session.closed`, retain the latest observed usage and mark final usage as unconfirmed.579An HTTP session-creation error means the session did not reach `session.started`. Handle startup errors separately from errors in a running session. If a running connection fails before `session.closed`, retain the latest observed usage and mark final usage as unconfirmed.
548 580
549If a stored session is available, [fork it](#store-and-fork-a-session) to start a new session from its saved state. Otherwise, create a replacement session with relevant saved history. Reconcile pending actions with your backend before retrying them, and suppress stale results from the previous session. Restore application state explicitly rather than assuming a new connection resumes the previous session or its pending work.581If a completed stored recording is available, [fork it](#store-and-fork-a-session) to continue in a new session. Otherwise, create a replacement session with relevant saved history. Before continuing, check unfinished actions with your backend, restore current task state, and update result routing so late results from the previous session cannot overwrite newer work.