SpyBara
Go Premium

Documentation 2026-09-18 22:59 UTC to 2026-09-19 23:00 UTC

17 files changed +381 −243. View all changes and history on the product overview
2026
Tue 29 22:57 Mon 28 22:57 Sat 26 23:59 Fri 25 23:58 Thu 24 23:58 Wed 23 23:58 Tue 22 23:57 Mon 21 23:00 Sat 19 23:00 Fri 18 22:59 Thu 17 10:04 Wed 16 20:58 Tue 15 22:59 Mon 14 22:58 Sun 13 15:02 Fri 11 20:00 Thu 10 18:01 Wed 9 23:59 Sat 5 17:01 Fri 4 23:59 Thu 3 23:00 Wed 2 22:59
Details

15- `packages`: Install Python, system, or global `npm` packages with `python`, `system`, or `npm` lists. Pin versions when needed, such as `pandas==2.2.3`.15- `packages`: Install Python, system, or global `npm` packages with `python`, `system`, or `npm` lists. Pin versions when needed, such as `pandas==2.2.3`.

16- `setup_commands`: Run ordered shell commands before the agent starts, such as `[{ "command": "mkdir -p reports" }]`. Each command has its own optional `cwd`, defaulting to `/workspace`.16- `setup_commands`: Run ordered shell commands before the agent starts, such as `[{ "command": "mkdir -p reports" }]`. Each command has its own optional `cwd`, defaulting to `/workspace`.

17- `files`: [Supply input files](https://developers.openai.com/api/docs/guides/agents-api/environments/files#upload-files) by Files API ID or inline base64 content.17- `files`: [Supply input files](https://developers.openai.com/api/docs/guides/agents-api/environments/files#upload-files) by Files API ID or inline base64 content.

18- `env`: Set string-valued environment variables. Runtime-reserved names, including `PATH`, `CODEX_*`, and `OPENAI_API_KEY`, are rejected.18- `env`: Set string-valued environment variables. Agent-generated code can read these values. IMPORTANT: For secrets, use [vault credentials](https://developers.openai.com/api/docs/guides/agents-api/tools/vaults#use-vault-secrets-for-api-requests-from-a-sandbox) to keep the real values outside the sandbox. Runtime-reserved names, including `PATH`, `CODEX_*`, and `OPENAI_API_KEY`, are rejected.

19- `skills`, `plugins`, `capability_directories`: Add [skills](https://developers.openai.com/api/docs/guides/tools-skills#agents-api) and [plugins](https://developers.openai.com/api/docs/guides/agents-api/tools/plugins).19- `skills`, `plugins`, `capability_directories`: Add [skills](https://developers.openai.com/api/docs/guides/tools-skills#agents-api) and [plugins](https://developers.openai.com/api/docs/guides/agents-api/tools/plugins).

20- `environment_template_id`: [Reuse saved configuration](https://developers.openai.com/api/docs/guides/agents-api/tools/plugins#reuse-a-hosted-plugin-setup) across sessions. Omitted settings inherit the template; network overrides cannot broaden its policy.20- `environment_template_id`: [Reuse saved configuration](https://developers.openai.com/api/docs/guides/agents-api/tools/plugins#reuse-a-hosted-plugin-setup) across sessions. Omitted settings inherit the template; network overrides cannot broaden its policy.

21 21 

Details

41 41 

42## Broker third-party access42## Broker third-party access

43 43 

44Keep third-party credentials outside the environment. Where possible, route requests through a credential broker. The broker injects secrets into approved outbound requests without placing them in the agent's environment.44Keep third-party credentials outside the environment. For API requests from an OpenAI-hosted sandbox, use [vault secrets as environment variables](https://developers.openai.com/api/docs/guides/agents-api/tools/vaults#use-vault-secrets-for-api-requests-from-a-sandbox). Sandbox code uses a placeholder; a network proxy supplies the real secret for approved hosts.

45 

46For self-hosted environments, configure a trusted proxy or server to supply secrets outside the environment. This is infrastructure you provide. For [function tools](https://developers.openai.com/api/docs/guides/agents-api/tools/functions), keep credentials in the application that handles the call and return only the result.

45 47 

46<picture>48<picture>

47 <source49 <source

Details

160For a server that allows anonymous access, omit authentication fields and `vault_ids`. Otherwise, choose the credential source for your connection:160For a server that allows anonymous access, omit authentication fields and `vault_ids`. Otherwise, choose the credential source for your connection:

161 161 

162- **HTTP credentials for one session:** Set `transport.authorization` or `transport.headers` when creating the session. The Agents API encrypts these values and omits them from the returned session resource.162- **HTTP credentials for one session:** Set `transport.authorization` or `transport.headers` when creating the session. The Agents API encrypts these values and omits them from the returned session resource.

163- **Reusable HTTP credentials:** Store credentials in a [vault](https://developers.openai.com/api/docs/guides/agents-api/tools/vaults) and attach it through `vault_ids`. Vaults apply only to connections from OpenAI. Credentials match the server URL; use `credential_id` to select one when several match.163- **Reusable HTTP credentials:** Store MCP credentials in a [vault](https://developers.openai.com/api/docs/guides/agents-api/tools/vaults) and attach it through `vault_ids`. Vault-backed MCP authentication applies only to connections from OpenAI. Credentials match the server URL; use `credential_id` to select one when several match.

164- **Stdio credentials:** Supply values in the environment and list their names in `transport.env_vars`. These values can be read by code running in the environment. Self-hosted sessions do not accept inline values in `transport.env`.164- **Stdio credentials:** Supply values in the environment and list their names in `transport.env_vars`. These values can be read by code running in the environment. Self-hosted sessions do not accept inline values in `transport.env`.

165 165 

166For example, an HTTP transport can include a bearer token and another header:166For example, an HTTP transport can include a bearer token and another header:

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5A vault stores credentials for MCP connections from OpenAI. Attach it to a session so the agent can use authenticated tools without receiving the secret values.5A vault stores credentials outside your agent's instructions and configuration. Attach it to a session with `vault_ids` so the session can use those credentials.

6 6 

7Vaults support bearer tokens and existing OAuth grants. For connections from your environment, use the other [MCP authentication options](https://developers.openai.com/api/docs/guides/agents-api/tools/mcp#add-authentication).7Choose the credential type based on where the request runs:

8 

9| Request | Credential type | How the session uses the credential |

10| ----------------------------------------- | ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |

11| MCP connection from OpenAI | `static_bearer` or `mcp_oauth` | OpenAI authenticates to the configured MCP server. |

12| API request from an OpenAI-hosted sandbox | `environment_variable` | Code uses an environment variable containing a placeholder. A network proxy replaces the placeholder with the secret for approved hosts. |

13 

14For example, a sandbox can use a vault secret to call the GitHub REST API. Follow [Use vault secrets for API requests from a sandbox](#use-vault-secrets-for-api-requests-from-a-sandbox).

15 

16Retrieving a vault or credential does not return its secret values. For other MCP connections, see [MCP authentication options](https://developers.openai.com/api/docs/guides/agents-api/tools/mcp#add-authentication).

8 17 

9## Permissions18## Permissions

10 19 


16 25 

17 26 

18 27 

19## Create and use a vault28 

29 

30 

31 

32## Create and use a vault for MCP secrets

20 33 

21Use your API client, the MCP server URL (`mcp_url`), and an access token for that server (`access_token`). The examples use GitHub tools.34Use your API client, the MCP server URL (`mcp_url`), and an access token for that server (`access_token`). The examples use GitHub tools.

22 35 


336```349```

337 350 

338 351 

339The Agents API selects a credential that matches the server URL. If several attached credentials match, set the MCP tool's `credential_id` to select one. Retrieving a vault or credential does not return its secret values.352The Agents API selects a credential that matches the server URL. If several attached credentials match, set the MCP tool's `credential_id` to select one.

353 

354 

355 

356 

357## Use vault secrets for API requests from a sandbox

358 

359Use an `environment_variable` credential to supply a secret for API requests from an OpenAI-hosted sandbox. The sandbox receives a placeholder in the named environment variable. The real secret stays outside the sandbox.

360 

361This workflow requires an `openai_hosted` environment. It does not supply credentials to self-hosted environments or application-run [function tools](https://developers.openai.com/api/docs/guides/agents-api/tools/functions).

362 

363### Store the API token

364 

365Create a vault as shown above. Then send a credential creation request to `POST /v1/vaults/{vault_id}/credentials` with these fields:

366 

367| Field | Value for a GitHub API token |

368| ------------------------------- | ----------------------------------------------------------------------- |

369| `name` | `GitHub API token` |

370| `auth.type` | `environment_variable` |

371| `auth.secret_name` | `GITHUB_TOKEN` |

372| `auth.secret_value` | The token, read from a secret environment variable in your application. |

373| `auth.networking.type` | `limited` |

374| `auth.networking.allowed_hosts` | `["api.github.com"]` |

375 

376`secret_name` is the environment variable that sandbox code reads. `secret_value` is the real credential. Keep it out of prompts, source files, and logs.

377 

378Use exact host names in `allowed_hosts`, without a scheme, path, port, or wildcard. The proxy supplies secrets only to HTTPS destinations on port 443 or 8443.

379 

380### Attach the vault to a hosted session

381 

382Include the following fields alongside `agent` when [creating a session](https://developers.openai.com/api/docs/guides/agents-api/sessions#create-a-session). Replace `vault_123` with the vault ID returned by the API:

383 

384```json

385{

386 "vault_ids": ["vault_123"],

387 "environment": {

388 "type": "openai_hosted",

389 "network": {

390 "access": "restricted",

391 "allowed_domains": ["api.github.com"]

392 }

393 }

394}

395```

396 

397The two host lists control different things. `allowed_domains` lets the sandbox connect to a host. The credential's `allowed_hosts` lets the proxy supply the secret to that host.

398 

399With restricted network access, include every credential host in `allowed_domains`. Do not set `network.access` to `disabled` for a session with environment credentials.

400 

401Each attached environment credential must have a unique `secret_name`. Do not also define that name in `environment.env`.

402 

403### Call the API from the sandbox

404 

405[Send the agent a message](https://developers.openai.com/api/docs/guides/agents-api/sessions#send-input) asking it to run this command in the sandbox:

406 

407```bash

408curl https://api.github.com/user \

409 -H "Authorization: Bearer $GITHUB_TOKEN"

410```

411 

412The command reads the placeholder from `GITHUB_TOKEN`. The proxy replaces it with the real token before sending the request to `api.github.com`. A successful request returns the authenticated GitHub user's account details as JSON. Printing the variable inside the sandbox shows the placeholder, not the token.

413 

414Pass the placeholder unchanged in the HTTPS request header. It cannot supply the real secret for local computation, such as signing a request. For those tasks, keep the credential in your application and expose the operation through a [function tool](https://developers.openai.com/api/docs/guides/agents-api/tools/functions).

340 415 

341 416 

342 417 


526 601 

527## Rotate or remove credentials602## Rotate or remove credentials

528 603 

529[Update a credential](https://developers.openai.com/api/reference/resources/beta/subresources/agents/subresources/vaults/subresources/credentials/methods/update) to replace its token without changing its ID, authentication type, or server URL. For OAuth, use the saved `vault_id` and `credential_id` with the replacement token and expiry:604[Update a credential](https://developers.openai.com/api/reference/resources/beta/subresources/agents/subresources/vaults/subresources/credentials/methods/update) to replace its secret without changing its ID or authentication type. For MCP credentials, the server URL also stays the same. For OAuth, use the saved `vault_id` and `credential_id` with the replacement token and expiry:

530 605 

531Rotate an OAuth token606Rotate an OAuth token

532 607 


644 719 

645Include `expires_at` when the replacement token expires. Supplying a new access token without an expiry clears the stored expiry; an explicit `null` also clears it.720Include `expires_at` when the replacement token expires. Supplying a new access token without an expiry clears the stored expiry; an explicit `null` also clears it.

646 721 

722For an environment credential, send `auth.type: "environment_variable"` and the replacement `auth.secret_value` to `POST /v1/vaults/{vault_id}/credentials/{credential_id}`. Create a new session to use the replacement. Updating the vault does not change the credential already configured in an existing sandbox.

723 

724To change `secret_name` or `networking`, create a new credential.

725 

647[Delete a credential](https://developers.openai.com/api/reference/resources/beta/subresources/agents/subresources/vaults/subresources/credentials/methods/delete) when you no longer need it. [Delete a vault](https://developers.openai.com/api/reference/resources/beta/subresources/agents/subresources/vaults/methods/delete) to remove the vault and all its credentials.726[Delete a credential](https://developers.openai.com/api/reference/resources/beta/subresources/agents/subresources/vaults/subresources/credentials/methods/delete) when you no longer need it. [Delete a vault](https://developers.openai.com/api/reference/resources/beta/subresources/agents/subresources/vaults/methods/delete) to remove the vault and all its credentials.

648 727 

649Deleting stored credentials does not revoke the original tokens with their providers or stop a running session. Your application handles provider-side revocation and [session cancellation](https://developers.openai.com/api/docs/guides/agents-api/sessions#cancel-an-active-turn).728Deleting stored credentials does not revoke the original tokens with their providers or stop a running session. Your application handles provider-side revocation and [session cancellation](https://developers.openai.com/api/docs/guides/agents-api/sessions#cancel-an-active-turn).

Details

174sample needs at least five seconds of actual speech and at least 15 transcribed175sample needs at least five seconds of actual speech and at least 15 transcribed

175text tokens; silence does not count. Use a 10–30-second176text tokens; silence does not count. Use a 10–30-second

176recording with several complete sentences. Each upload is limited to 10 MiB.177recording with several complete sentences. Each upload is limited to 10 MiB.

177The service extracts the reference transcript; do not upload transcript tokens,178The service extracts the reference transcript from your recording.

178configure a decoder, or add custom request headers.

179 179 

180Browser recorders may label audio `audio/webm;codecs=opus`, which the upload180Browser recorders may label audio `audio/webm;codecs=opus`, which the upload

181endpoint rejects. When constructing an upload, use the supported base MIME type181endpoint rejects. When constructing an upload, use the supported base MIME type


184 184 

185### Select the voice at session creation185### Select the voice at session creation

186 186 

187Pass a custom voice as the object `{ "id": "voice_123" }`, not the string187Pass a custom voice as an object, such as `{ "id": "voice_123" }`. Named voices

188`"voice_123"`. Named voices such as `"marin"` use strings.188such as `"marin"` use strings.

189 189 

190`gpt-live-1` supports custom voices with English accents. To use an accent, also190`gpt-live-1` supports custom voices with English accents. To use an accent, also

191specify it in `session.instructions`, such as "Speak British English" or "Speak191specify it in `session.instructions`, such as "Speak British English" or "Speak


202}202}

203```203```

204 204 

205For [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live), the trusted session broker205For [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live), have your application server

206places this configuration in the JSON `session` field alongside `transport`.206place this configuration in the JSON `session` field alongside `transport`.

207The Live endpoint requires JSON, not multipart or raw SDP. Read the created207Authenticate requests to your server with application credentials, and keep the

208session ID from `session.id` and the SDP answer from `transport.sdp`. Authenticate208OpenAI API key on the server.

209hosted broker requests with application credentials; never expose the OpenAI

210API key to the browser.

211 209 

212For [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live), put the configuration210For [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live), put the configuration

213in the first `session.start` event. Connect without query parameters and wait211in the first `session.start` event. Follow the connection guide for streaming

214for `session.started` before streaming audio. Send audio with212audio and closing the session.

215`session.input_audio.append`. After sending `session.close`, keep receiving until

216`session.closed` supplies final usage.

217 213 

218### Handle access and lifecycle failures214### Handle access and lifecycle failures

219 215 

220- The output voice cannot be changed after the Live session starts. Start a new session to use a different voice.216- Choose the voice at session creation. Start a new session to use a different voice.

221- A deleted or revoked voice, a consent from another project, or missing custom-voice access can appear as a `404`.217- A deleted or revoked voice, a consent from another project, or missing custom-voice access can appear as a `404`.

222- Malformed audio, a mismatched speaker, or a non-project-scoped key is rejected.218- Malformed audio, a mismatched speaker, or a non-project-scoped key is rejected.

223 219 

guides/live.md +11 −10

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5GPT-Live handles a spoken conversation while a backend agent looks up information, uses tools, and completes tasks. It can listen while speaking, a capability called **full duplex**. Sending work to the backend is called **delegation**: the conversation can continue while that work runs.5GPT-Live handles a spoken conversation while a backend agent looks up information, uses tools, and completes tasks. It can listen while speaking (**full duplex**). Sending work to the backend is called **delegation**.

6 6 

7For example, a user can ask about an order, add a detail while the backend checks its status, and hear the result when it is ready. You choose the backend model or agent independently of the voice model; Realtime uses one model for speech, reasoning, and tool selection.7For example, a user can ask about an order and add a detail while the backend checks its status. GPT-Live can keep talking with the user and explain the result when it arrives. You choose the backend model or agent independently of the voice model.

8 8 

9## Understand the two parts9## Understand the two parts

10 10 

11- **GPT-Live handles conversation.** It listens, speaks, and decides when to ask the backend for help. Give it a short prompt for conversation style and when to delegate.11- **GPT-Live handles conversation.** It listens, speaks, and decides when to ask the backend for help. Give it a short prompt for conversation style and when to delegate.

12- **The backend handles delegated tasks.** With Responses delegation, use a supported Responses model. With client delegation, connect any model, agent harness, or service your application runs. The backend reasons, uses tools, and returns results for GPT-Live to communicate. Keep detailed instructions, business rules, and tool workflows here.12- **The backend handles delegated tasks.** With Responses delegation, use a supported Responses model. With client delegation, connect any model, agent harness, or service your application runs. The backend reasons, uses tools, and returns results for GPT-Live to communicate. Keep detailed instructions, business rules, and tool workflows here.

13 13 

14Your application owns permissions, confirmations, private function execution, and durable task state. Interrupting speech does not automatically cancel backend work. See [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) to compare GPT-Live with Realtime and chained voice applications.14Your application checks permissions, obtains required confirmations, runs functions that access your systems, and saves task progress. Backend work can continue when the caller interrupts the assistant; your application decides whether to finish or cancel it. See [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) to compare GPT-Live with Realtime and chained voice applications.

15 15 

16 16 

17 17 


19 19 

20## Choose how to run the backend20## Choose how to run the backend

21 21 

22Start with **[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#configure-responses-delegation)** when a managed backend fits: GPT-Live calls your configured Responses model, supplies conversation context, and returns results. Choose **[client delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#configure-client-delegation)** when your application needs to control backend execution, context, or which results reach GPT-Live.22Start with **[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#configure-responses-delegation)** to have OpenAI run the backend model and pass conversation context and results between it and GPT-Live. Your application still runs your own function tools. Choose **[client delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#configure-client-delegation)** to connect an existing agent or control the backend’s context, execution, and returned results yourself.

23 23 

24See [Choose a delegation mode](https://developers.openai.com/api/docs/guides/live-delegation#choose-a-delegation-mode) for the comparison and configuration details. Choose the mode when you create the session; to change modes, start a new session.24See [Choose a delegation mode](https://developers.openai.com/api/docs/guides/live-delegation#choose-a-delegation-mode) for the comparison and configuration details. Choose the mode when you create the session; to change modes, start a new session.

25 25 


343. Wait for `session.started`, then speak and listen to a reply. Ask a question that needs current information to try the web search backend.343. Wait for `session.started`, then speak and listen to a reply. Ask a question that needs current information to try the web search backend.

354. End the conversation and [close the session](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) to collect final usage and release the connection.354. End the conversation and [close the session](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) to collect final usage and release the connection.

36 36 

37Check both the spoken conversation and the backend result. A session-start event confirms startup; listening to a reply and checking the search result verify separate parts of the application.37For your first test, listen to the assistant and check that its answer reflects the backend’s search result.

38 38 

39GPT-Live voice sessions are billed by duration, per second. See the [model pricing](https://developers.openai.com/api/docs/models/gpt-live-1) for the current rate. Backend model and tool usage is billed separately. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting and ways to reduce costs.39GPT-Live voice sessions are billed by duration, per second. See the [model pricing](https://developers.openai.com/api/docs/models/gpt-live-1) for the current rate. Backend model and tool usage is billed separately. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting and ways to reduce costs.

40 40 

41## Choose a connection41## Choose a connection

42 42 

43- **[WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live)** for browser voice applications. Media tracks carry audio; a data channel carries JSON events.43- **[WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live)** for browser voice applications. It carries microphone and speaker audio on media tracks and JSON events on a data channel.

44- **[WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live)** for server-side audio integrations. The primary socket carries audio and control events.44- **[WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live)** for server-side audio integrations. One connection carries audio and control events.

45- **[Server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live)** for backend access to an existing session. A sideband connection carries events while audio stays on the primary connection.45- **[Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip?api=live)** for connecting phone calls.

46- **[Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip?api=live)** for phone integration paths and provider guidance.46 

47To monitor or control an existing session from your backend, add a [server-side connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live). This additional WebSocket is called a **sideband**; audio continues through the session’s primary connection.

47 48 

48## Partner integrations49## Partner integrations

49 50 

50For applications built with **LiveKit**, **Twilio**, **Telnyx**, or **Daily/Pipecat**, start with the [partner integration overview](https://developers.openai.com/api/docs/guides/live-partner-integrations) to choose a connection for your existing media path.51If your application already uses **LiveKit**, **Twilio**, **Telnyx**, or **Daily/Pipecat**, follow the [partner integration overview](https://developers.openai.com/api/docs/guides/live-partner-integrations) to connect its existing calls or audio streams to GPT-Live.

51 52 

52## Continue building53## Continue building

53 54 

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5After [connecting to GPT-Live](https://developers.openai.com/api/docs/guides/live), use session events to update context, display transcripts, and manage the connection's lifecycle. The model can listen and speak at the same time, so keep received events, audio playback, and backend task state separate in your application.5After [connecting to GPT-Live](https://developers.openai.com/api/docs/guides/live), use session events to add context, display transcripts, and manage the connection. GPT-Live can listen and speak at the same time. Track transcript text, played audio, and backend task progress separately so your interface can show what the assistant is saying and what work is still running.

6 6 

7This guide assumes your connection has emitted `session.started`. See [Connections](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for connection setup and audio streaming, and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for backend work.7This guide assumes your connection has emitted `session.started`. See [Connections](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for connection setup and audio streaming, and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for backend work.

8 8 


20| ------------ | --------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |20| ------------ | --------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |

21| Model | Set the required `model`. | Start a new session to change it. |21| Model | Set the required `model`. | Start a new session to change it. |

22| Instructions | Set `instructions` for conversation behavior, up to 16,384 tokens. | Add instructions with `session.instructions.append`. |22| Instructions | Set `instructions` for conversation behavior, up to 16,384 tokens. | Add instructions with `session.instructions.append`. |

23| History | Set `input` to relevant prior text messages. It defaults to `[]`. | Append context; don't replace the startup history. |23| History | Set `input` to relevant prior text messages. It defaults to `[]`. | Add context with append events. |

24| Voice | Set `audio.output.voice` to a supported voice or authorized custom voice. The default is `marin`. | Start a new session to change it. |24| Voice | Set `audio.output.voice` to a supported voice or authorized custom voice. The default is `marin`. | Start a new session to change it. |

25| Delegation | Set `delegation.type` to `client` or `responses`. Omitted or `null` delegation selects client mode. | Update Responses settings within the existing mode. |25| Delegation | Set `delegation.type` to `client` or `responses`. Omitted or `null` delegation selects client mode. | Update Responses settings within the existing mode. |

26| Storage | Set `store` to `true` to make the session available for forking. It defaults to `false`. | Choose at startup. |26| Storage | Set `store` to `true` to make the session available for forking. It defaults to `false`. | Choose at startup. |


44| Delta | `delta` | English | Southern U.S. | Feminine | Generated |44| Delta | `delta` | English | Southern U.S. | Feminine | Generated |

45| Cinder | `cinder` | English | Southern U.S. | Masculine | Generated |45| Cinder | `cinder` | English | Southern U.S. | Masculine | Generated |

46 46 

47Regional influence describes a voice's speaking style, not a guarantee of accent fidelity. For an approved voice created from your own recording, see [Custom voices](https://developers.openai.com/api/docs/guides/custom-voices).47Regional influence describes a voice’s speaking style. Test the voice with the languages and pronunciation your application needs. For an approved voice created from your own recording, see [Custom voices](https://developers.openai.com/api/docs/guides/custom-voices).

48 48 

49 49 

50 50 

51 51 

52 52 

53For WebSocket, choose the shared `audio.format` at startup; it cannot change during the session. For WebRTC, omit this field because the connection negotiates its audio format. See [WebSocket audio formats](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for format and streaming details.53For WebSocket, choose `audio.format` at startup. The same format applies to input and output audio for the session. To use another format, start a new session. WebRTC negotiates its audio format during connection setup, so leave `audio.format` out of WebRTC requests. See [WebSocket audio formats](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for supported formats and streaming details.

54 54 

55### Update a live session55### Update a live session

56 56 

57Use `session.update` for changes to `session.delegation.responses` in a session already using Responses delegation. Send only the settings you want to change; omitted settings retain their values. See [Configure Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for the settings and update workflow.57Use `session.update` for changes to `session.delegation.responses` in a session already using Responses delegation. Send only the settings you want to change; omitted settings retain their values. See [Configure Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for the settings and update workflow.

58 58 

59You cannot change the delegation mode after startup. In particular, setting `delegation` to `null` selects client mode; it does not reset a Responses session. The startup fields `model`, `instructions`, `input`, `audio`, and `store` are not accepted update fields. Unknown configuration fields are rejected.59Choose the delegation mode and the fields `model`, `instructions`, `input`, `audio`, and `store` at startup. Use `session.update` only for the supported Responses settings described above; other configuration fields are rejected. To switch delegation modes, create a new session. At startup, `delegation: null` selects client delegation rather than restoring default Responses settings.

60 60 

61A successful update emits `session.updated` with the full resolved session configuration. When you supply an `event_id`, the acknowledgment returns it as `client_event_id`. Check for [rejected commands](#handle-rejected-commands) as well as acknowledgments. Acceptance confirms the configuration update; it does not establish that a backend task ran or that the model spoke.61A successful update emits `session.updated` with the resulting session configuration. Match it to your outgoing `event_id` through `client_event_id`. Handle [rejected commands](#handle-rejected-commands) in the same event loop. Track backend work and spoken output through their own events.

62 62 

63## Provide history and context63## Provide history and context

64 64 


94```94```

95 95 

96 96 

97The list accepts up to 128 messages and 8,192 combined tokens. Supported roles are `developer`, `user`, and `assistant`, each with one text part. Developer and user messages use `input_text`; assistant messages use `text` or `output_text`. Put trusted application instructions in `instructions` or a developer message. The list does not accept the `system` role.97The list accepts up to 128 messages and 8,192 combined tokens. Each message has one text part and one of these roles: `developer`, `user`, or `assistant`. Developer and user messages use `input_text`; assistant messages use `text` or `output_text`. Put trusted application instructions in `instructions` or a developer message.

98 98 

99Select the history needed for the next interaction. `input` is a startup field, not a way to replace history during a running session. It also does not accept the full range of backend input items used in Responses delegation.99Select the text history needed for the next interaction and supply it at startup. During the session, add updates with the context events below. Send backend-specific items, such as tool results, through the [delegation workflow](https://developers.openai.com/api/docs/guides/live-delegation).

100 100 

101### Understand when context reaches the model101### Understand when context reaches the model

102 102 

103The full `input` supplied at session creation is available to the model when the session starts. Put context the model needs from the beginning in this field.103Put any context the model needs from the start in `input`; the full field is available when the session starts.

104 104 

105During a running session, the `session.instructions.append`, `session.thinking.append`, and `session.commentary.append` events feed content into the model over time. Their acknowledgments wait until frame progress reaches the estimated end of context injection. The returned `start_ms` and `end_ms` describe an estimated range on the session timeline, not speech or playback completion. They do not prove that the model consumed the entire update. Don't assume its next speech will reflect the whole update.105During a running session, `session.instructions.append`, `session.thinking.append`, and `session.commentary.append` add context over time. The acknowledgment arrives when the session timeline reaches the estimated end of the added context. Its `start_ms` and `end_ms` estimate where that update falls on the session timeline.

106 106 

107If frame progress stops, an acknowledgment can remain pending. Closing the session reports an error for pending appends. Match each acknowledgment to the outgoing `event_id` through `client_event_id`, and keep handling errors while you wait.107These times describe context delivery, not speech or playback. The model may still respond before it has used the whole update. When an action depends on a new instruction or fact, verify the resulting behavior in your application.

108 

109If the session timeline stops, the acknowledgment can remain pending. Match acknowledgments to the outgoing `event_id` through `client_event_id`, and keep handling errors while you wait. Closing the session returns errors for appends that are still pending.

108 110 

109### Add context during the conversation111### Add context during the conversation

110 112 


147```149```

148 150 

149 151 

150Wait for `session.thinking.appended` with `client_event_id: "context_1"`, or handle an error. The acknowledgment confirms that context was accepted. It does not confirm speech, playback, or completion of an external action.152Handle `session.thinking.appended` with `client_event_id: "context_1"`, or the corresponding error, to track this update. See [Understand when context reaches the model](#understand-when-context-reaches-the-model) for acknowledgment timing.

151 153 

152Quiet context can influence later speech; it is not a privacy boundary. Keep credentials, secrets, and text the model must never reveal out of all three events. Use the instructions event for application-authored behavior, not untrusted tool output. Enforce permissions and required confirmations in your application.154The assistant may repeat information supplied through any of these events. Send only information suitable for the conversation, and keep credentials and secrets in your backend. Use `session.instructions.append` for behavior defined by your application. Supply factual tool results as context, and enforce permissions and required confirmations in application code.

153 155 

154For page navigation, selections, and other UI changes, see [Share UI context](https://developers.openai.com/api/docs/guides/live-delegation#share-ui-context) for concise updates that help GPT-Live understand what the user is referring to.156For page navigation, selections, and other UI changes, see [Share UI context](https://developers.openai.com/api/docs/guides/live-delegation#share-ui-context) for concise updates that help GPT-Live understand what the user is referring to.

155 157 

156For results tied to a backend task, use a known client delegation ID and follow [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update). That ID is not a Responses response ID or tool call ID.158For an update about a specific backend task, use the ID of the relevant client delegation. A delegation ID identifies the Live task; Responses response IDs and tool call IDs identify different objects. See [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update) for the workflow.

159 

160When your application detects a problem, send a short correction through the session’s primary WebSocket or a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband). See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for checks, action controls, and playback handling.

161 

162 

157 163 

158Use instructions to steer the conversation after an application check triggers. Your server can monitor events and send these corrections through a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) attached to the existing session, or through its primary WebSocket. See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for concurrent checks, action blocking, and playback control.

159 164 

160 165 

161 166 


164 169 

165 170 

166 171 

167## Build the conversation interface

168 172 

169Display transcripts and microphone state independently of backend progress. Receiving assistant text does not tell you how much audio the user has heard.173## Manage speech and transcripts

170 174 

171### Transcript deltas175### Transcript deltas

172 176 


182}186}

183```187```

184 188 

185Append fragments in order for each speaker, retaining `start_ms` and `end_ms`. These are milliseconds on the session timeline, with intervals that include the start and exclude the end. They are not wall-clock timestamps, packet arrival times, or exact word alignments.189Append each speaker’s `delta` fragments exactly as received, preserving spaces and repeated words. Retain their `start_ms` and `end_ms`. These values are milliseconds from the start of the session. The example above covers the interval from 1,000 ms up to, but excluding, 1,200 ms. They describe approximate fragment timing rather than exact word boundaries; use them instead of packet arrival times to organize the transcript.

186 190 

187Only intervals containing transcript text produce events, and network delivery can be uneven. Do not infer silence from a missing event or treat a fragment as a complete user turn. Transcript deltas have no item ID or authoritative turn-completed event.191Transcript events arrive for intervals that contain text, and delivery can be uneven. A fragment may contain only part of a sentence; a gap in delivery may be a network delay. Transcript deltas have no item ID or event that marks a completed conversational turn, so your application decides how to group them for display.

188 192 

189Processing transcript fragments is optional. You can use them to update your UI, run checks, or start work early while the conversation continues. For lightweight checks, consider a small model such as `gpt-5.6-luna` with low reasoning effort. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples and connection guidance.193Processing transcript fragments is optional. You can use them to update your UI, run checks, or start work early while the conversation continues. For lightweight checks, consider a small model such as `gpt-5.6-luna` with low reasoning effort. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples and connection guidance.

190 194 

191For conversation guardrails, check accumulated user and assistant text as it arrives. Transcript delivery does not provide an advance buffer for approving speech before playback. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).195Use [transcript guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#run-checks-alongside-the-conversation) to monitor the conversation and trigger interventions while speech continues. If your application needs to check assistant speech before playback, see [Check speech before playback](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#check-speech-before-playback) for buffering, approval, interruption, and recovery handling.

192 196 

193 197 

194 198 

195 199 

196 200 

197If your interface groups text into turns, keep that grouping revisable. Preserve the original fragments, allow user and assistant intervals to overlap, and tune any gap timeout against recorded conversations. A brief acknowledgment from the other speaker may belong within an ongoing exchange. Grouping fragments must not trigger tool execution or cancel backend work by itself.201Keep transcript timing separate from audio playback. WebSocket `session.output_audio.delta` events have no timing fields or output-audio-done event; WebRTC delivers audio through its media track. See [Connections](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for audio handling.

198 

199 202 

200 203 

201 204 

202 205 

203Keep transcript timing separate from audio playback. WebSocket `session.output_audio.delta` events have no timing fields or output-audio-done event; WebRTC delivers audio through its media track. See [Connections](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for audio handling.

204 206 

205### Display captions207### Display captions

206 208 

207Build caption rows that can grow while both speakers are talking:209GPT-Live is full duplex: the caller and assistant can speak at the same time. Update their captions independently so both speakers’ text can keep growing during overlapping speech.

208 210 

2091. **Preserve the text.** Store each speaker's original `delta`, `start_ms`, and `end_ms`. Concatenate text exactly as received, including spaces and repeated words. Don't trim fragments or insert spaces between them.211If your app uses chat bubbles, the fragments “I’d like” and “ to change my booking” can appear in one caller bubble. If the assistant says “Sure” while the caller continues, show that acknowledgment separately while allowing the caller’s bubble to keep growing. Keep the original fragments and timestamps so text that arrives later can update the appropriate bubble.

2102. **Update each speaker independently.** Allow user and assistant rows to grow during overlapping speech. Keep earlier assistant text visible after an interruption, and start a new row when the assistant resumes.

2113. **Keep rows stable.** Assign display IDs in your application and preserve row order as text grows. Don't derive row identity from changing text or end timestamps, or move a row to the bottom whenever it receives a fragment.

2124. **Revisit grouping for late fragments.** Use transcript timestamps to group nearby fragments from the same speaker. Allow late text to update earlier rows and revise fragment assignments while retaining the original fragments. These display groups are not complete semantic turns; any gap threshold is an application choice to test.

2135. **Let the reader control scrolling.** Follow new text while the reader is at the bottom. Pause automatic scrolling when they scroll up, and provide a way to return to the latest captions.

2146. **Show tool progress in a status area.** Use assistant transcript events for spoken captions. Display tool activity and backend results outside the captions; receiving a result does not mean the assistant has said it.

215 212 

216Test the display with overlapping speech, short acknowledgments, interruptions, long pauses, and translation where the two speakers' text arrives at different rates.213Use `session.output_transcript.delta` for spoken captions and show backend updates separately. Keep decisions about running tools or canceling work in your application’s task logic, separate from how you group text for display.

217 214 

218### Control microphone input215### Control microphone input

219 216 


244 241 

245Wait for `session.input_audio.muted` with `client_event_id: "mute_1"` before treating the command as accepted. To resume input, send `session.input_audio.unmute` and wait for `session.input_audio.unmuted`. Handle errors for either command.242Wait for `session.input_audio.muted` with `client_event_id: "mute_1"` before treating the command as accepted. To resume input, send `session.input_audio.unmute` and wait for `session.input_audio.unmuted`. Handle errors for either command.

246 243 

247Muting input does not stop inference, delegated work, or generated speech. Control microphone capture and audio playback separately in your application when those controls are needed.244Muting input leaves the session running: the model can keep generating speech, and delegated work can continue. Use your application’s microphone capture and audio player controls when you also need to stop local recording or playback.

248 245 

249### Greet before the caller speaks246### Greet before the caller speaks

250 247 

251To request a greeting after `session.started`:248To have GPT-Live open the conversation, send greeting instructions after `session.started`. Specify the language, what the assistant should say, and that it should begin immediately, then pause to listen. Use the application’s chosen greeting language until the caller speaks. For example:

252 249 

2531. Send one fresh `session.instructions.append` with `delegation_id: null`. Include the greeting, its language, and an explicit instruction to greet immediately without waiting for the caller, then pause and listen. Keep the existing startup instructions.250> Greet the caller now in English. Introduce yourself as the support assistant and ask how you can help. Then pause and listen.

2542. Wait for `session.instructions.appended`, matching its `client_event_id` to your command. Handle a rejected command before continuing.

2553. Keep input audio running, including silence before the caller speaks. On WebSocket, continue sending `session.input_audio.append`; on WebRTC, keep the negotiated input audio track active. Observe output transcript and audio for the greeting.

256 251 

257Use the language specified by your application until the caller speaks; don't infer it from a name, phone number, or location. See [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for prompt design.2521. Keep input audio running throughout this sequence, including silence before the caller speaks. On WebSocket, continue sending `session.input_audio.append`; on WebRTC, keep the input audio track active.

2532. Send the instructions once with `session.instructions.append` and `delegation_id: null`.

2543. Match `session.instructions.appended` to your command using `client_event_id`, and handle any error. This acknowledgment confirms that the instructions were accepted.

258 255 

259For a greeting that needs to follow application instructions, send those instructions with `session.instructions.append`, then use a short `session.commentary.append` to prompt the assistant to begin. For example: “Begin the conversation now, following the instructions provided.” Keep input audio running, including silence before the caller speaks.256For exact wording and a known playback-completion point, play a verified recording or rendered clip through your application and [control GPT-Live playback while it plays](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed). Test greetings in the languages you support, including when the caller starts speaking during the greeting. See [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for prompt design.

260 

261Instructions request a greeting; they do not guarantee exact wording or uninterrupted playback. The API does not emit an opening-completed event, and acknowledgment does not mean the greeting was heard. Use application-controlled playback if the audio must be verbatim. Test your greeting with the languages and interruptions your application supports.

262 257 

263### Deliver a disclosure258### Deliver a disclosure

264 259 


296```291```

297 292 

298 293 

299Keep input audio running as described in [Greet before the caller speaks](#greet-before-the-caller-speaks). Choose the delivery point deliberately: an instruction sent during the conversation can interrupt speech in progress.294Keep input audio running, as in [Greet before the caller speaks](#greet-before-the-caller-speaks). An instruction sent during the conversation can interrupt speech in progress.

300 295 

301This requests the wording; it does not guarantee exact delivery. Verify the complete spoken disclosure and actual playback before marking it delivered. `session.instructions.appended` confirms only that the instruction was accepted. If exact audio delivery is required, play a verified recording or rendered clip through your application and control GPT-Live output while it plays. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).296Check the generated disclosure and its playback before marking it delivered. The instruction acknowledgment records acceptance; use the audio itself to check the wording. For exact wording and a known playback-completion point, play a verified recording or rendered clip through your application and control GPT-Live output while it plays. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).

302 297 

303 298 

304 299 


306 301 

307## Manage longer conversations302## Manage longer conversations

308 303 

309GPT-Live automatically manages context during long conversations; no configuration parameter is needed. The instructions you provide at session start are preserved throughout compaction. You don’t need to resend them.304GPT-Live manages long conversations automatically and preserves your original startup instructions.

310 305 

311The default context window holds 128,000 tokens, including your instructions, conversation text, and audio tokens that don’t appear in the transcript.306The default context window holds 128,000 tokens, including your instructions, conversation text, and audio tokens that don’t appear in the transcript.

312 307 


320 315 

321## Store and fork a session316## Store and fork a session

322 317 

323Set `store` to `true` in the session configuration at creation to save a recording for later download or forking. Storage defaults to `false` and must be enabled for your project. Downloads and forks require a completed stored recording and a data policy that permits persistence. Recordings expire after 30 days. With Zero Data Retention, `store` is treated as `false`, and recording downloads and forks are unavailable. See [GPT-Live data controls](https://developers.openai.com/api/docs/guides/your-data#v1livesessions).318A fork starts a new session from a saved voice conversation. Use it to run several evaluation trials from the same reference conversation, or to let a user continue after an earlier session has ended. Each fork gets a new connection and session ID, with the source conversation and its saved configuration as its starting point.

319 

320### Run evaluations from a reference conversation

321 

322Suppose you want to test how your agent handles a caller changing an order. Record the setup once, through the point where the caller has identified the order. End and finalize that session before the caller asks to change it. Each evaluation can then fork the same reference session and receive the same next caller audio: “Actually, can you send it to my office instead?”

323 

324For each trial, restore the same test order and application state, supply the next caller input, and evaluate the new response and tool actions. You can repeat the scenario or compare supported Responses backend settings. Measure fork startup separately from response time. See the [GPT-Live evaluation guide](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) for choosing scenarios and measuring results.

325 

326Forks inherit the GPT-Live model, voice, and original instructions. To compare a different voice model or startup prompt, create new sessions with that configuration. The fork API uses the completed source recording, so end the reference session where you want the evaluation to begin.

327 

328### Continue after a session ends

324 329 

325For example, add this field to the `session` object in your WebSocket `session.start` event or WebRTC creation request:330For example, a caller may hang up and call back later, or reconnect after a dropped call. If the earlier session has a completed stored recording, your application can fork it on a new connection and continue from the saved conversation.

331 

332Save application task state alongside the source session ID. Before continuing, check the status of any outstanding backend work and give the new session its current results. For example, if an order update was already submitted, confirm its outcome before attempting another update. Use the new session ID for controls and sideband connections, and route subsequent backend results to the new session.

333 

334### Prepare a session for forking

335 

3361. **Enable storage when you create the source.** Set `store: true` in its session configuration. Storage defaults to `false`, must be enabled for your project, and requires a data policy that permits persistence.

3372. **Save the source session ID.** Read it from `session.started` or the WebRTC creation response, and associate it with your application’s conversation record.

3383. **Finish and close the source.** Complete required backend work, then follow [Usage and graceful close](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close). Keep the connection open until `session.closed` and handle any finalization error. Forking requires a completed stored recording; saving it can add time to finalization.

3394. **Start a fork on a new connection.** Use the source ID with the transport flow below, save the new session ID, and complete startup before continuing the conversation. Set the child’s `store` explicitly: `true` if you want to fork its continuation later, or `false` if you do not need to store that trial. Omitting it inherits the source setting.

340 

341Stored recordings are available for 30 days. With Zero Data Retention (ZDR), `store` is treated as `false` and forking is unavailable. For fork-based evaluations, use a non-ZDR organization with storage enabled for the project.

342 

343If you have no completed stored recording, [start a new session with relevant saved text history](https://developers.openai.com/api/docs/guides/live-conversations#seed-a-session-with-prior-conversation). See [GPT-Live data controls](https://developers.openai.com/api/docs/guides/your-data#v1livesessions) for storage requirements.

344 

345For example, set this field in the source session’s WebSocket `session.start` configuration or WebRTC creation request:

326 346 

327```json347```json

328{348{


330}350}

331```351```

332 352 

333Save the source session ID from `session.started` or the WebRTC creation response. A fork starts a **new session with a new ID** from the stored session state. It does not reopen the original connection or reuse the source session ID.

334 

335Start the fork through the transport your application uses:353Start the fork through the transport your application uses:

336 354 

337| Transport | Start the fork |355| Transport | Start the fork |


339| WebSocket | Connect to `wss://api.openai.com/v1/live/sessions/{source_session_id}/fork`. |357| WebSocket | Connect to `wss://api.openai.com/v1/live/sessions/{source_session_id}/fork`. |

340| WebRTC | Send a new SDP offer to `POST /v1/live/sessions/{source_session_id}/fork`. Apply the returned `transport.sdp` answer to the new peer connection. |358| WebRTC | Send a new SDP offer to `POST /v1/live/sessions/{source_session_id}/fork`. Apply the returned `transport.sdp` answer to the new peer connection. |

341 359 

342A fork inherits the stored session configuration, subject to the transport rules below. For a WebSocket fork, send `session.start` with a required `session` object; `{}` supplies no overrides. Do not supply a new model or repeat the original instructions or input. You can override `store`, Responses delegation settings, and the new WebSocket audio format. WebRTC forks can override `store`, Responses delegation settings, and frontend client permissions. Omitting `store` on a fork inherits the source session's setting.360A fork inherits the source session’s model, original instructions, and input. Send only the supported overrides at startup:

361 

362- **WebSocket:** `store`, Responses delegation settings, and the new connection’s `audio.format`. Send a `session.start` event with a `session` object; use `{}` to keep inherited settings where supported.

363- **WebRTC:** `store`, Responses delegation settings, and frontend client permissions.

343 364 

344A WebSocket fork does **not** inherit the source audio format: set `audio.format` explicitly or use the default PCM16 at 24 kHz. It also discards inherited frontend data-channel permissions. WebRTC forks negotiate their audio format and reject `audio.format`; they preserve frontend permission settings unless you override them.365For a WebSocket fork, set `audio.format` for the new connection or use the default PCM16 at 24 kHz. The source audio format and frontend data-channel permissions are not inherited. WebRTC negotiates audio format during connection setup; omit `audio.format`. WebRTC preserves frontend permission settings unless you override them.

345 366 

346Wait for `session.started` before sending further WebSocket commands. WebRTC starts through the HTTP request and must not receive a second `session.start` on its data channel.367For WebSocket, wait for `session.started` before sending more commands. For WebRTC, the HTTP request starts the session; continue through the negotiated connection without sending another `session.start`.

347 368 

348### Start a WebSocket fork369### Start a WebSocket fork

349 370 


438 459 

439Return the response to your frontend, apply `transport.sdp` as the new peer connection's answer, and retain the new `session.id`. Keep the API key on your backend.460Return the response to your frontend, apply `transport.sdp` as the new peer connection's answer, and retain the new `session.id`. Keep the API key on your backend.

440 461 

441Use the new session ID for later sideband connections and session controls. Keep application task state separately: restoring conversation state does not confirm that a pending backend action completed. Reconcile uncertain results before retrying an action. If you don't have a stored session to fork, [seed a new session with saved history](#seed-a-session-with-prior-conversation).462Use the new session ID for sideband connections and session controls. Before retrying an unfinished action, check its outcome in your backend and restore the current application task state. If you have no completed stored recording, [seed a new session with saved history](#seed-a-session-with-prior-conversation).

442 463 

443### Download a recording464### Download a recording

444 465 


470```491```

471 492 

472 493 

494## Close idle sessions and resume

495 

496For applications with long gaps between interactions, close the voice session during inactivity and start a new session when the user returns. Keep conversation context and application task state so the user can continue without repeating themselves. For example, an in-car assistant can resume when the driver activates voice again, while a coding assistant can keep its backend worker running between voice conversations.

497 

4981. **Decide when to close.** Use an application-controlled inactivity timeout based on audio activity, assistant playback, and application interactions. Allow for expected pauses, such as reading or thinking. Close only when playback has finished and no pending work requires the current voice session. Gaps between transcript events alone do not establish silence.

4992. **Save state and close gracefully.** Save the source session ID, conversation context, and current task state. Finish any required Responses work, then follow [Usage and graceful close](#usage-and-graceful-close): install the `session.closed` listener, send `session.close`, and wait for `session.closed` before releasing the connection. With client delegation, application-managed backend work can continue independently while voice is closed.

5003. **Detect when to restart.** Offer a button labeled **Resume conversation**, a push-to-talk control, or an application-managed wake trigger. A closed Live session cannot listen for the user. If you use local speech detection to restart automatically, keep microphone capture active and buffer the opening speech through connection setup. Deliver that audio once the new session is ready, so the user’s first words are preserved.

5014. **Restore context in a new session.** If the source was created with `store: true`, storage is enabled and permitted, and its recording finalized successfully, [fork the stored session](#prepare-a-session-for-forking). Otherwise, [start a new session with saved text history](#seed-a-session-with-prior-conversation). Save the new session ID, check the status of outstanding backend operations, and route subsequent results to the new session. Keep operation status in your application so restarting does not repeat completed actions.

502 

503Muting the microphone leaves the session active. Choose an idle timeout by comparing avoided voice duration with session-creation costs and the delay before voice becomes ready again. See [Voice session costs](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live#voice-session-costs) and [WebRTC initialization charges](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live#webrtc-initialization-charges).

504 

473## Handle errors and end the session505## Handle errors and end the session

474 506 

475Keep reading session events until the session finalizes. Distinguish a rejected command, a failed connection, and a completed session so your application can recover appropriately.507Keep reading session events until the session finalizes. Distinguish a rejected command, a failed connection, and a completed session so your application can recover appropriately.


492}524}

493```525```

494 526 

495An error code can be `null`, and an error may lack a client event ID. Handle those cases without assuming a command succeeded. For an immutable-field error, keep the current configuration or create a new session with the intended settings.527Provide a general error handler for errors whose code is `null` or whose client event ID is absent. For an immutable-field error, keep the current configuration or create a new session with the intended settings.

496 528 

497### Handle moderation529### Handle moderation

498 530 


501- Some moderation events end the session.533- Some moderation events end the session.

502- Others cut off assistant audio for the remainder of its current speech and emit an `error` event without ending the session.534- Others cut off assistant audio for the remainder of its current speech and emit an `error` event without ending the session.

503 535 

504Read `error` events even while audio is playing. Don't assume every moderation error closes the session, or that an audio interruption means the connection failed. Keep application state aligned with the session lifecycle, and don't mark an interrupted spoken message as fully delivered. Application-level [conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) remain separate from this built-in moderation behavior.536Keep handling `error` events while audio is playing. Track audio interruption and session closure separately, and mark a spoken message as delivered only after checking its playback. Apply your own [conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) alongside built-in moderation.

505 537 

506### Usage and graceful close538### Usage and graceful close

507 539 


516}548}

517```549```

518 550 

519These are snapshots, not increments to sum. Backend token usage is separate; preserve it from nested Responses completion events. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting.551Use the latest `usage.seconds` as the running total for voice duration. For example, updates of 12 and then 15 seconds mean 15 seconds of use. Track backend token usage separately from nested Responses completion events. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting.

520 552 

521To close gracefully:553To close gracefully:

522 554 


528 560 

529Sending `session.close` cancels queued Responses and rejects further commands. An active response can finish, but one waiting for a function result cannot continue after closing starts. Decide separately whether to finish or cancel work your application runs through client delegation.561Sending `session.close` cancels queued Responses and rejects further commands. An active response can finish, but one waiting for a function result cannot continue after closing starts. Decide separately whether to finish or cancel work your application runs through client delegation.

530 562 

531The `session.closed` event establishes finalization; the embedded session is a configuration snapshot. A socket close alone does not establish success, and a transport close code after a valid final event does not invalidate finalization. Closing WebRTC immediately after sending the command can prevent delivery of the final event.563Use `session.closed` to confirm finalization and read the final configuration snapshot. Keep the transport open until this event arrives. If the socket closes first, record finalization as unconfirmed; if it closes after a valid `session.closed`, retain the confirmed result.

532 564 

533The final event's `reason` explains why the session ended:565The final event's `reason` explains why the session ended:

534 566 


546 578 

547An HTTP session-creation error means the session did not reach `session.started`. Handle startup errors separately from errors in a running session. If a running connection fails before `session.closed`, retain the latest observed usage and mark final usage as unconfirmed.579An HTTP session-creation error means the session did not reach `session.started`. Handle startup errors separately from errors in a running session. If a running connection fails before `session.closed`, retain the latest observed usage and mark final usage as unconfirmed.

548 580 

549If a stored session is available, [fork it](#store-and-fork-a-session) to start a new session from its saved state. Otherwise, create a replacement session with relevant saved history. Reconcile pending actions with your backend before retrying them, and suppress stale results from the previous session. Restore application state explicitly rather than assuming a new connection resumes the previous session or its pending work.581If a completed stored recording is available, [fork it](#store-and-fork-a-session) to continue in a new session. Otherwise, create a replacement session with relevant saved history. Before continuing, check unfinished actions with your backend, restore current task state, and update result routing so late results from the previous session cannot overwrite newer work.

Details

6 6 

7Read more about [steering the live model for delegation and tools](https://developers.openai.com/api/docs/guides/live-prompting#delegation) in the prompting guide.7Read more about [steering the live model for delegation and tools](https://developers.openai.com/api/docs/guides/live-prompting#delegation) in the prompting guide.

8 8 

9The event examples on this page use `connection`, a connected primary Live WebSocket or sideband from the [connection guides](https://developers.openai.com/api/docs/guides/voice-websockets?api=live). On a primary connection, wait for `session.started` before calling an event helper. An attached sideband already belongs to a running session.

10 

9 11 

10 12 

11 13 


14 16 

15With **[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#configure-responses-delegation)**, GPT-Live calls the Responses model you choose, supplies conversation context, and returns backend results to the live conversation. With **[client delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#configure-client-delegation)**, your application prepares the context, runs an agent or workflow, and sends results back to GPT-Live.17With **[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#configure-responses-delegation)**, GPT-Live calls the Responses model you choose, supplies conversation context, and returns backend results to the live conversation. With **[client delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#configure-client-delegation)**, your application prepares the context, runs an agent or workflow, and sends results back to GPT-Live.

16 18 

17Start with Responses delegation when its managed workflow fits. Choose client delegation when you need more control over backend context, execution, or the results returned to GPT-Live.19Start with Responses delegation if you want GPT-Live to manage requests. Choose client delegation when you need to run your own workflow or review results before sending them to GPT-Live.

18 20 

19 21 

20 22 


32 34 

33For example, a travel assistant can send flight-status questions to an airline service and itinerary changes to a separate planning agent. The application chooses which backend to call and what verified result to return to GPT-Live.35For example, a travel assistant can send flight-status questions to an airline service and itinerary changes to a separate planning agent. The application chooses which backend to call and what verified result to return to GPT-Live.

34 36 

35In both modes, your application manages task state and enforces permissions and required confirmations before running its custom tools. Reviewing backend results is a separate decision: it does not approve every word GPT-Live speaks or guarantee silence while validation runs. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).37In both modes, your application tracks task progress and checks permissions and required user confirmations before running custom tools. GPT-Live can continue speaking while your application reviews a backend result. If your application must control when the user hears audio, add [playback controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).

36 38 

37Client delegation also requires your application to [maintain conversation context](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#keep-the-conversation-context-in-your-application). The delegation event contains metadata, not task text; use transcript events and application state to prepare the backend request.39Client delegation also requires your application to [maintain conversation context](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#keep-the-conversation-context-in-your-application). The delegation event contains metadata, not task text; use transcript events and application state to prepare the backend request.

38 40 


65 67 

66Start with [GPT-5.6 Terra](https://developers.openai.com/api/docs/models/gpt-5.6-terra), or try [GPT-5.6 Luna](https://developers.openai.com/api/docs/models/gpt-5.6-luna) for cost-sensitive workloads. Compare answer quality and latency on your tasks before choosing a backend model.68Start with [GPT-5.6 Terra](https://developers.openai.com/api/docs/models/gpt-5.6-terra), or try [GPT-5.6 Luna](https://developers.openai.com/api/docs/models/gpt-5.6-luna) for cost-sensitive workloads. Compare answer quality and latency on your tasks before choosing a backend model.

67 69 

68Register supported tools in `delegation.responses.tools`. Use `delegation.responses.tool_choice` to control which tools the backend can use: `"auto"` lets it choose, `"required"` requires a tool call, and `"none"` prevents one. You can also select a named function. Set `delegation.responses.parallel_tool_calls` to `true` to allow independent lookups together, or `false` when calls must run sequentially. Your application still executes its custom functions and enforces dependencies and approvals. These settings do not force the live model to delegate.70Register supported tools in `delegation.responses.tools`. Set `delegation.responses.tool_choice` to `"auto"` to let the backend choose a tool, `"required"` to require a tool call, or `"none"` to disable tool calls. You can also select a named function.

71 

72Set `delegation.responses.parallel_tool_calls` to `true` to allow multiple tool calls in a response, or `false` for sequential calls. Your application executes custom functions and checks their dependencies and required approvals. These settings apply after GPT-Live delegates; use the [live prompt](https://developers.openai.com/api/docs/guides/live-prompting#delegation) to guide when it should delegate.

69 73 

70The Responses configuration requires a backend `model` at creation. It supports `function` definitions and `web_search` entries in `tools`. It also exposes `max_output_tokens` (at least 16 when set), `service_tier`, and the `reasoning` and `text` settings supported by the selected backend model. See [Reduce backend latency](#reduce-backend-latency) for settings you can tune.74The Responses configuration requires a backend `model` at creation. It supports `function` definitions and `web_search` entries in `tools`. It also exposes `max_output_tokens` (at least 16 when set), `service_tier`, and the `reasoning` and `text` settings supported by the selected backend model. See [Reduce backend latency](#reduce-backend-latency) for settings you can tune.

71 75 

72If [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode) is available for your model and project, consider it for latency-sensitive calls. For GPT-Live, select it with `delegation.responses.service_tier: "priority"`.76If [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode) is available for your model and project, consider it for latency-sensitive calls. For GPT-Live, select it with `delegation.responses.service_tier: "priority"`.

73 77 

74As the conversation changes, send `session.update` with changes in `session.delegation.responses` to update the backend model, instructions, available tools, `tool_choice`, or other supported settings without starting a new Live session. Omitted settings retain their values. Setting `delegation` to `null` selects client mode and cannot reset a running Responses session; switching modes fails with `immutable_field_update`.78Send `session.update` with changes in `session.delegation.responses` to update the backend model, instructions, tools, `tool_choice`, or other supported settings during the conversation. Omitted settings keep their current values.

79 

80To switch between Responses and client delegation, create a new Live session. Updating `delegation` to `null` selects client mode, so sending it to a running Responses session fails with `immutable_field_update`.

75 81 

76These settings use familiar Responses concepts, but Live supports a subset of the standalone Responses API. Live supplies conversation context and initiates delegated work. Configure the backend through the session; the Live `response.create` command uses that configuration and does not accept a standalone Responses request body.82These settings use familiar Responses concepts, but Live supports a subset of the standalone Responses API. Live supplies conversation context and initiates delegated work. Configure the backend through the session; the Live `response.create` command uses that configuration and does not accept a standalone Responses request body.

77 83 


100}106}

101```107```

102 108 

103Dispatch on `envelope.event.type` and preserve the outer `delegation_id`. Do not handle every top-level `response.*` value as an unwrapped Responses event. Tolerate additional nested Responses lifecycle events.109When the top-level event type is `response.event`, dispatch on `envelope.event.type`. Save the outer `delegation_id` to associate the backend event with its delegation. Your handler may also receive nested Responses lifecycle events beyond those shown here.

104 110 

105Live speech and delegated work continue independently. A completed backend response does not itself mean the user heard the answer. Use the Live output transcript and audio for the spoken part of the interaction.111Live speech and delegated work continue independently. A completed backend response does not itself mean the user heard the answer. Use the Live output transcript and audio for the spoken part of the interaction.

106 112 

107### Complete a client-actionable function call113 

114 

115 

116 

117### Run a custom function and return its result

108 118 

109Read completed function calls from nested `response.output_item.done` events. The finished function item contains `call_id`, `name`, and `arguments`; an arguments-done event alone is not sufficient to identify the call.119Read completed function calls from nested `response.output_item.done` events. The finished function item contains `call_id`, `name`, and `arguments`; an arguments-done event alone is not sufficient to identify the call.

110 120 

111Track the response ID from nested `response.created` alongside the outer `delegation_id`, and collect that response's function calls from `response.output_item.done`. Forwarded lifecycle snapshots deliberately contain `response.output: []`, including at `response.completed`; their `tools` array is empty, `instructions` is `null`, and `input` is omitted. An empty terminal output list does **not** mean there are no pending function calls. Use the collected calls to determine which results must be submitted before continuing.121Track the response ID from nested `response.created` alongside the outer `delegation_id`. Collect the response's function calls from `response.output_item.done`, and use that collection to determine which tool results to submit before continuing.

122 

123The forwarded lifecycle events, including `response.completed`, contain `response.output: []` even when function calls need results. These events also have an empty `tools` array, `instructions: null`, and no `input` field. Read the individual output-item events for the function calls.

112 124 

113After executing the authorized operation, append the result as a Responses item:125After executing the authorized operation, append the result as a Responses item:

114 126 


172```184```

173 185 

174 186 

175Submit every required result for the pending tool calls before continuing. Appending a function result does not automatically continue the response. `response.item.create` has no standalone success acknowledgment; keep processing errors and the subsequent nested response lifecycle.187Send a `response.item.create` result for every pending function call, then send `response.create` to continue the backend response. `response.item.create` has no separate success acknowledgment; keep processing errors and nested Responses lifecycle events.

176 188 

177`response.create` is a Live command for creating or continuing delegated Responses work, using the session's configured backend. Do not attach a Responses API creation body, backend model override, or `delegation_id` to this event. Both commands require Responses delegation.189Both commands require Responses delegation. The Live `response.create` command uses the backend configuration stored in the session. Use the event payload shown above, and configure the backend model and other settings through the session.

178 190 

179 191

180 192 


195```207```

196 208 

197 209 

198This selects client delegation for the session. Configure the backend separately: your application chooses its model or service, instructions, tools, and how to route work. If you use the Responses API for that backend, set its model and tools in your own Responses requests. The Live session does not configure or run those backend tools.210Your application configures and runs the backend: choose its model or service, instructions, tools, and routing. If the backend uses the Responses API, set its model and tools in your application's Responses requests.

199 211 

200When GPT-Live requests help, your application builds the backend request from conversation and application context, runs the work, and decides which results to send back. Enforce permissions and required confirmations before executing your tools. Retain the full conversation history in your application so you can provide the relevant context for each backend request.212When GPT-Live requests help, build the backend request from your saved conversation history and current task state. Check permissions and required confirmations, run the work, and choose which results to return.

201 213 

202## Keep the conversation context in your application214## Keep the conversation context in your application

203 215 


226}238}

227```239```

228 240 

229Read `event.delegation.id`. The delegation object contains metadata, not task text. Maintain the transcript and application context needed by your own delegated-work handler. Current IDs have an `item_` prefix, as illustrated here; treat the full ID as opaque and return it unchanged rather than constructing or parsing one.241Save `event.delegation.id` and include it unchanged in updates about this task.

230 242 

231Return a result using that ID:243Return a result using that ID:

232 244 


257```269```

258 270 

259 271 

260Use `session.thinking.append` to add information to the model's internal reasoning without speaking it aloud when appended. Use `session.commentary.append` for a result the model should speak aloud; the model is trained to paraphrase the appended text. All appends contain a plain string and require `delegation_id`, including when its value is `null`. A non-null ID must name a known client delegation.272Use `session.commentary.append` for results GPT-Live should say aloud; it is trained to paraphrase the text. Use `session.thinking.append` for facts or progress that it can use in later replies without saying them when they arrive. You can send multiple updates with the same client delegation ID.

261 273 

262Repeated result appends can continue the same client delegation. An appended acknowledgment arrives after estimated context injection; it is not proof that the model has consumed or spoken the result, or that an external action succeeded.274See [Send the right kind of update](#send-the-right-kind-of-update) for content limits, required fields, and acknowledgment timing.

263 275 

264 276

265 277 


282and confirmation requirements.]294and confirmation requirements.]

283 295 

284## Return the result296## Return the result

285Return the relevant facts, whether the task is complete, and what comes next.297Return the relevant facts, the task's current status, and the next step.

286Use confirmed values. Do not invent a successful action.298Report an action as complete after the tool or service confirms success.

299If the outcome is unclear, state that and explain what needs to be checked.

287```300```

288 301 

289Keep large structured payloads, lengthy tool output, and Markdown intended for display in the backend. Give GPT-Live the relevant facts and let it choose how to say them. A concise tool result doesn't need an additional model call to rewrite it for speech.302Keep large structured payloads, lengthy tool output, and Markdown intended for display in the backend. Give GPT-Live the relevant facts and let it choose how to say them. A concise tool result doesn't need an additional model call to rewrite it for speech.

290 303 

291With client delegation, [return the result directly to GPT-Live](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#receive-a-client-delegation). With Responses delegation, follow the [function-result flow](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#complete-a-client-actionable-function-call) to continue backend work.304With client delegation, [return the result directly to GPT-Live](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#receive-a-client-delegation). With Responses delegation, follow the [function-result flow](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#complete-a-client-actionable-function-call) to continue backend work.

292 305 

293SDK event examples below use `connection`, a connected primary Live WebSocket or sideband from the [connection guides](https://developers.openai.com/api/docs/guides/voice-websockets?api=live). Call the helper after `session.started` on a primary connection; an attached sideband already belongs to a running session.

294 

295## Send the right kind of update306## Send the right kind of update

296 307 

297Choose an event based on how GPT-Live should use the content:308Choose an event based on how GPT-Live should use the content:


403 414 

404The corresponding acknowledgements are `session.thinking.appended`, `session.commentary.appended`, and `session.instructions.appended`. Match their `client_event_id` to your outgoing `event_id`. The acknowledgment waits for estimated context injection, not for speech or playback to finish. See [when context reaches the model](https://developers.openai.com/api/docs/guides/live-conversations#understand-when-context-reaches-the-model) for timing and error handling.415The corresponding acknowledgements are `session.thinking.appended`, `session.commentary.appended`, and `session.instructions.appended`. Match their `client_event_id` to your outgoing `event_id`. The acknowledgment waits for estimated context injection, not for speech or playback to finish. See [when context reaches the model](https://developers.openai.com/api/docs/guides/live-conversations#understand-when-context-reaches-the-model) for timing and error handling.

405 416 

406Quiet context can still affect what the model says later. It is not a private place for secrets or hidden reasoning. Send useful facts and brief progress summaries.417Send facts and brief progress summaries that GPT-Live can use in the conversation. Content sent with `session.thinking.append` can influence later spoken replies; keep secrets and private backend reasoning in your application.

407 418 

408## Keep updates accurate and useful419## Keep updates accurate and useful

409 420 


420| Failed | “That time is no longer available.” |431| Failed | “That time is no longer available.” |

421| Cancellation confirmed | “Your appointment has been canceled.” |432| Cancellation confirmed | “Your appointment has been canceled.” |

422 433 

423A spoken interruption does not automatically cancel backend work. If the user changes Friday to Thursday, update the active task and ignore late Friday results. Your application must decide whether to cancel work, change it, or let it finish. Check that cancellation succeeded before saying it did.434When the user changes a request, update the active task in your application. For example, if they change Friday to Thursday, use Thursday for subsequent work and ignore results from the outdated Friday request. Manage any cancellation in the backend, and confirm it succeeded before telling the user that the work was canceled. Interrupting the spoken conversation leaves backend work running.

424 435 

425Before retrying a failed tool call, check whether the original action already happened. For example, a lost response should not cause a second booking. If the outcome is unclear, say so and offer the next useful step.436Before retrying a failed tool call, check whether the original action already happened. For example, a lost response should not cause a second booking. If the outcome is unclear, say so and offer the next useful step.

426 437 


437 448 

438### Accept typed input449### Accept typed input

439 450 

440If a caller types an exact value, such as an order number, pass it to the backend that handles the task. A voice-only application does not need this path. Keep the typed value as user data rather than a live-model instruction.451Send typed values, such as order numbers, to the backend as user-provided data so it can use the exact text.

441 452 

442 453 

443 454 


536 547 

537### Responses delegation548### Responses delegation

538 549 

539Live manages persistent WebSocket connections to Responses, prepares the connection and known request configuration in advance, and reuses prior response state when available. You do not need to implement those steps for the hosted backend. Reuse depends on the active connection and compatible state; it does not guarantee a cache hit or a specific latency.550Live manages persistent WebSocket connections to Responses and prepares known request configuration in advance. It can also reuse prior response state when the active connection and state support it. Measure response time and reported cache usage on your workload to see the effect.

540 551 

541Tune the backend through `delegation.responses`:552Tune the backend through `delegation.responses`:

542 553 


559- **Prepare known configuration.** Initialize instructions, tools, and connections before the first request needs them. Responses WebSocket mode also supports warming up known request state before generation; follow its [setup guidance](https://developers.openai.com/api/docs/guides/websocket-mode#connect-and-create-responses).570- **Prepare known configuration.** Initialize instructions, tools, and connections before the first request needs them. Responses WebSocket mode also supports warming up known request state before generation; follow its [setup guidance](https://developers.openai.com/api/docs/guides/websocket-mode#connect-and-create-responses).

560- **Stream useful results.** Return coherent, verified chunks with `session.commentary.append`. Use `session.thinking.append` for quiet progress. Preserve the client delegation ID and the 500-token limit per append. Keep private reasoning in the backend and confirm actions before announcing success.571- **Stream useful results.** Return coherent, verified chunks with `session.commentary.append`. Use `session.thinking.append` for quiet progress. Preserve the client delegation ID and the 500-token limit per append. Keep private reasoning in the backend and confirm actions before announcing success.

561- **Keep reusable input stable.** Preserve instructions, tool definitions and ordering, and unchanged history prefixes. Append new information after reusable content when your backend supports caching and continuation.572- **Keep reusable input stable.** Preserve instructions, tool definitions and ordering, and unchanged history prefixes. Append new information after reusable content when your backend supports caching and continuation.

562- **Avoid unnecessary buffering.** Forward a useful result as soon as it is ready. Buffer only enough to classify the output and form a coherent chunk. Prefer structured phase metadata; if you use text prefixes to distinguish progress from results, wait for the complete prefix before forwarding text.573- **Forward complete, useful updates.** Send each result as soon as you have enough text to understand it on its own. Have your backend label updates as progress or results so your application can choose the appropriate append event. If it marks these categories with a text prefix, wait for the full prefix before forwarding the update.

563 574 

564Measure the first useful spoken answer when comparing this path with Responses delegation.575Measure the first useful spoken answer when comparing this path with Responses delegation.

565 576 


583 594 

584Process accumulated text when meaningful new information arrives. A fragment may be incomplete, and later speech can change the request. Discard outdated results, coordinate with subsequent delegated work to avoid duplicate actions, and apply your usual permission and confirmation checks before consequential actions.595Process accumulated text when meaningful new information arrives. A fragment may be incomplete, and later speech can change the request. Discard outdated results, coordinate with subsequent delegated work to avoid duplicate actions, and apply your usual permission and confirmation checks before consequential actions.

585 596 

586To feed information back into the conversation:597Send findings or instructions back to GPT-Live using the [append event that matches the update](#send-the-right-kind-of-update). For work started outside a client delegation, use `delegation_id: null`. Your application applies UI changes and manages tool execution and cancellation.

587 

588| Intent | Event |

589| ------------------------------------------------------------- | ----------------------------- |

590| Change the live model's behavior or redirect the conversation | `session.instructions.append` |

591| Provide quiet context for subsequent responses | `session.thinking.append` |

592| Provide information the model should say aloud | `session.commentary.append` |

593 

594For updates outside a client delegation, use `delegation_id: null`. These appends steer the live model; your application controls UI changes, tool execution, and cancellation. See [Send the right kind of update](#send-the-right-kind-of-update) for append examples.

595 598 

596### Shared optimizations599### Shared optimizations

597 600 


606 609 

607## Verify the complete interaction610## Verify the complete interaction

608 611 

609Test both the authoritative application state and the audio the client played. A backend response can finish while the spoken result is interrupted, and a context acknowledgment confirms acceptance rather than playback. Keep operation IDs and task revisions separate from delegation IDs so reconnects, retries, and late results do not repeat or reverse an action.612Verify that the backend completed the intended action and that the client played the expected spoken result. For example, check both the booking record and the audio played after a successful reservation. The backend can finish while the spoken answer is interrupted, so test these outcomes separately.

613 

614Track each application action with its own operation ID and each changed request with a task revision, alongside the GPT-Live delegation ID. Use those records to recognize completed work after reconnects or retries and to discard results for outdated requests.

610 615 

611Use [Evaluating voice agents](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) for repeatable tests. For an existing Realtime tool loop or chained backend, follow [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration).616Use [Evaluating voice agents](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) for repeatable tests. For an existing Realtime tool loop or chained backend, follow [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration).

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5GPT-Live handles the voice conversation while a backend handles task reasoning and tools. Keep your application logic, tool implementations, permissions, and durable state. The migration connects those responsibilities to the new voice interface.5GPT-Live handles listening and speaking. A backend decides how to complete tasks and which tools to call. Keep your existing tool implementations, permission checks, and saved task records. During migration, connect that backend to GPT-Live and decide which instructions belong in each model.

6 6 

7This guide uses an appointment assistant: check availability, ask the user to confirm a slot, then book it. Start with a connected session from [Getting started](https://developers.openai.com/api/docs/guides/live), and keep representative conversations from your existing application for comparison.7This guide uses an appointment assistant: check availability, ask the user to confirm a slot, then book it. Start with a connected session from [Getting started](https://developers.openai.com/api/docs/guides/live), and keep representative conversations from your existing application for comparison.

8 8 


22 22 

23## Choose your delegation mode23## Choose your delegation mode

24 24 

25Your existing architecture is a useful starting point:25Choose who will run the backend:

26 26 

27- **Responses delegation** fits a Realtime app where the model selects functions and your application executes them. A hosted Responses model takes over task reasoning and tool selection.27- **Responses delegation:** Configure a hosted Responses model to reason about tasks and select tools. Your application executes custom functions and returns their results. This is a useful starting point when your Realtime model currently selects those functions.

28- **Client delegation** fits an existing text agent or orchestrator. Your application supplies context, invokes that backend, and returns results to GPT-Live.28- **Client delegation:** Keep your existing agent or orchestrator. Your application supplies its conversation context, starts its work, and decides which results to send to GPT-Live.

29 29 

30Either migration path can use either mode. For example, a Realtime app that already has a separate backend agent may keep it with client delegation. Also consider how much control you need over backend context, execution, and reviewing results before they reach GPT-Live. See [Choose a delegation mode](https://developers.openai.com/api/docs/guides/live-delegation#choose-a-delegation-mode) for the full comparison.30Either mode can support either migration path. For example, a Realtime application with a separate backend agent can keep that agent through client delegation. See [Choose a delegation mode](https://developers.openai.com/api/docs/guides/live-delegation#choose-a-delegation-mode) for the full comparison.

31 31 

32## Choose your migration path32## Choose your migration path

33 33 

34Select the path that matches the application you have today.34Start with [From Realtime API](https://developers.openai.com/api/docs/guides/live-migration?migration-path=realtime#from-realtime-api) if your current voice model selects tools. Start with [From a text agent or chained pipeline](https://developers.openai.com/api/docs/guides/live-migration?migration-path=text-agent#from-a-text-agent-or-chained-pipeline) if you are keeping an existing agent and adding GPT-Live as its voice interface.

35 35 

36 36 

37 37 


543. Your application runs the function, returns its result, and continues the backend response.543. Your application runs the function, returns its result, and continues the backend response.

554. GPT-Live uses the answer from the backend to discuss available slots with the user.554. GPT-Live uses the answer from the backend to discuss available slots with the user.

56 56 

57GPT-Live can keep the conversation going while backend work runs. Finishing that work does not mean the assistant has finished speaking. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for configuration and the full event flow.57GPT-Live can continue speaking while the backend works. Track the backend task and audio playback separately: use tool results to update task status and your player’s state to update the speaking indicator. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for configuration and the full event flow.

58 

59### Preserve decisions that depend on audio

60 

61Check whether existing tool decisions depend on acoustic evidence, such as a voicemail beep or a recorded greeting's cadence. GPT-Live hears the incoming audio, but its voice frontend delegates work instead of issuing ordinary structured function calls. In client mode, `session.delegation.created` carries metadata and timing, without raw audio, task text, or parsed tool arguments. A delegated backend does not automatically receive the waveform.

62 

63For answering-machine detection, explicitly route incoming audio to an audio-capable detector. One application-managed architecture to evaluate runs a separate Realtime session alongside GPT-Live for part of the call:

64 

651. Send a copy of the incoming call audio to both sessions.

662. Have the detector report its classification through a structured function call. Check each result against your schema, reject stale results, and keep an unknown state when evidence is insufficient. Allow later evidence to revise the decision.

673. Send relevant trusted context to GPT-Live, and apply your application's policy to outgoing audio playback.

68 

69Keep human or machine classification separate from recording readiness. Recognizing voicemail does not establish that the greeting and beep have finished or that recording can begin. A classifier result or context acknowledgment also does not establish permission to play audio. Use [Adapt your guardrails](#adapt-your-guardrails) and the [playback controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed) to enforce that decision in the audio path your application controls.

70 

71Test a short “hello” that develops into a voicemail greeting, call-screening prompts, and a person picking up during voicemail. If you stop the detector before the call ends, test later human pickup after it stops. Choose when to stop the detector based on these tests and its added cost. The first human classification alone does not establish that further detection is unnecessary.

72 58 

73### Adapt the connection and audio lifecycle59### Adapt the connection and audio lifecycle

74 60 


85| Display user captions from input transcription events. | Append `session.input_transcript.delta` text to the user's captions. |71| Display user captions from input transcription events. | Append `session.input_transcript.delta` text to the user's captions. |

86| Display assistant captions from `response.output_audio_transcript.delta`. | Append `session.output_transcript.delta` text to the assistant's captions. |72| Display assistant captions from `response.output_audio_transcript.delta`. | Append `session.output_transcript.delta` text to the assistant's captions. |

87 73 

88**Generation and playback:** In Realtime, `response.output_audio.done` marks the end of audio generation, while `response.done` marks the end of the response stream. These events can also occur when a response is interrupted or unsuccessful; check `response.status` in `response.done`. Neither confirms that buffered audio has finished playing. For example, the server can finish generating while the client still has a second of audio to play. Drive a "speaking" indicator from playback state.74**Generation and playback:** Drive the speaking indicator from your audio player. The server can finish generating while the player still has a second of audio queued. In Realtime, `response.output_audio.done` marks the end of generation and `response.done` ends the response stream; check `response.status` for interruption or failure. GPT-Live has no equivalent event for the end of each spoken response.

89 75 

90**Captions:** Input transcription represents the user's speech; output transcription represents the assistant's generated speech. When input transcription is enabled, Realtime sends updates through `conversation.item.input_audio_transcription.delta` and a final transcript through `conversation.item.input_audio_transcription.completed`. A `delta` is a new text fragment. In GPT-Live, append each fragment to the corresponding speaker's captions independently because listening and speaking can overlap. A fragment is not a complete turn or confirmation of playback. See [Display captions](https://developers.openai.com/api/docs/guides/live-conversations#display-captions) for a display recipe.76**Captions:** When input transcription is enabled, Realtime sends text fragments through `conversation.item.input_audio_transcription.delta` and a final transcript through `conversation.item.input_audio_transcription.completed`. In GPT-Live, append each fragment to the caller’s or assistant’s captions; both can change at once. Your application decides how to group text and tracks playback through the audio player. See [Display captions](https://developers.openai.com/api/docs/guides/live-conversations#display-captions).

91 77 

92In GPT-Live, `response.create` starts or continues delegated Responses work. It does not grant permission for the voice model to speak. For startup, greetings, interruptions, and closing a session, follow [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).78Use `response.create` to start or continue delegated Responses work. GPT-Live manages when to speak as it listens to the conversation. For startup, greetings, interruptions, and closing a session, follow [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).

93 79 

94### Split conversation and backend instructions80### Split conversation and backend instructions

95 81 


134| Return each function result. | Send `conversation.item.create`. | Send `response.item.create`. |120| Return each function result. | Send `conversation.item.create`. | Send `response.item.create`. |

135| Continue after all required results. | Send `response.create`. | Send `response.create` to continue backend work. |121| Continue after all required results. | Send `response.create`. | Send `response.create` to continue backend work. |

136 122 

137For example, after `check_availability` returns one verified slot, your result changes as follows. These are messages on an already-connected session; `call_availability` stands for the actual call ID you received.123For example, after `check_availability` returns a verified slot, send the following on your connected session. Replace `call_availability` with the call ID you received.

138 124 

139**Before: Realtime result**125**Before: Realtime result**

140 126 


211```197```

212 198 

213 199 

214For the initial migration, setting `parallel_tool_calls` to `false` simplifies result handling. Collect calls from completed output-item events even if a terminal lifecycle snapshot has `output: []`. An arguments-done event alone does not supply the function name and `call_id`. Follow the complete [function-result procedure](https://developers.openai.com/api/docs/guides/live-delegation#complete-a-client-actionable-function-call) for collection, output submission, and errors.200For the initial migration, set `parallel_tool_calls` to `false` to handle one tool call at a time. Collect each function call from the inner `response.output_item.done` event and keep its name, arguments, and `call_id`. Keep that record even if a later completion event contains `output: []`. Wait for the completed item before running the handler; the arguments-done event alone lacks the function name and `call_id`. Follow the complete [function-result procedure](https://developers.openai.com/api/docs/guides/live-delegation#complete-a-client-actionable-function-call) for collection, output submission, and errors.

215 201 

216### Preserve context and apply corrections202### Preserve context and apply corrections

217 203 

218Responses delegation supplies relevant voice conversation context to the backend. Keep the authoritative appointment state in your application: selected slot, confirmed slot, permissions, active operation, and outcome. Live conversation history can be compacted; it is not your booking record.204Responses delegation supplies relevant voice conversation context to the backend. Keep the authoritative appointment state in your application: selected slot, confirmed slot, permissions, active operation, and outcome. Live conversation history can be compacted; it is not your booking record.

219 205 

220If the user says “Actually, Friday instead” while a Thursday lookup is pending, update the task's revision and invalidate the earlier slot confirmation. Before executing a booking, check that its arguments still match the current task and confirmation. Return an accurate superseded or cancelled result for any pending function call your application declines, then complete the required output batch before continuing. If a booking already succeeded, reconcile that result and the requested change before taking another action.206When the user says “Actually, Friday instead,” record Friday as the current request and clear any confirmation for Thursday. Give the task a new version number, such as revision 2, so your application can recognize results from the earlier request.

207 

208Before booking, check that the date, slot, and confirmation still match the current request. If you decline a pending function call because the user changed the request, return a result that explains it was skipped or cancelled, matching what actually happened. Submit a result for every required call before continuing the backend.

221 209 

222Transcript fragments can arrive late or overlap with assistant speech. Append each `delta` exactly as received and use `start_ms` and `end_ms` to group the display. These timestamps are not definitive turn boundaries or word-level playback timestamps. Clarify important dates, names, and numbers when intent is uncertain. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for transcript and context handling.210If the Thursday booking already succeeded, check its current status and handle the requested change before attempting another booking.

211 

212Keep each transcript fragment exactly as received, along with its speaker, `start_ms`, and `end_ms`. Use that information to update the appropriate caller or assistant caption or chat bubble, including when a fragment arrives late or both people speak at once. Choose message boundaries in your application and track audio playback in your player; the transcript timestamps do not identify exact word playback times. Clarify important dates, names, and numbers when intent is uncertain. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for transcript and context handling.

223 213 

224**Images and screen context:** If your Realtime application accepts images, route them to a vision-capable backend and return relevant text to GPT-Live. Both client and Responses delegation support this pattern. See [Add images and visual context](https://developers.openai.com/api/docs/guides/live-delegation#add-images-and-visual-context).214**Images and screen context:** If your Realtime application accepts images, route them to a vision-capable backend and return relevant text to GPT-Live. Both client and Responses delegation support this pattern. See [Add images and visual context](https://developers.openai.com/api/docs/guides/live-delegation#add-images-and-visual-context).

225 215 

216### Preserve decisions that depend on audio

217 

218Some decisions require the sound itself, such as detecting a voicemail beep or recognizing a recorded greeting from its timing. GPT-Live hears the call, but delegation does not automatically send audio to your backend. In client mode, `session.delegation.created` contains an ID and timing information; your application supplies the request context and any audio the backend needs.

219 

220For answering-machine detection, explicitly route incoming audio to an audio-capable detector. One application-managed architecture to evaluate runs a separate Realtime session alongside GPT-Live for part of the call:

221 

2221. Send a copy of the incoming call audio to both sessions.

2232. Have the detector report its classification through a structured function call. Check each result against your schema, reject stale results, and keep an unknown state when evidence is insufficient. Allow later evidence to revise the decision.

2243. Send relevant trusted context to GPT-Live, and apply your application's policy to outgoing audio playback.

225 

226Track two decisions: whether you are speaking to a person or a machine, and whether the destination is ready to record your message. A detector may recognize voicemail while the greeting is still playing. Wait for the evidence your application requires before allowing outgoing audio. A context acknowledgment records acceptance of the update; your application still makes the playback decision. Use [Adapt your guardrails](#adapt-your-guardrails) and the [playback controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed) to enforce that decision in the audio path your application controls.

227 

228Test a short “hello” that develops into a voicemail greeting, call-screening prompts, and a person picking up during voicemail. If you plan to stop the detector before the call ends, test what happens when a person picks up afterward. Use those results and the detector’s added cost to choose how long it should run. Include cases where an initial “human” classification changes as more audio arrives.

229 

226 230

227 231 

228 232 


251}255}

252```256```

253 257 

254The notification contains metadata, not request text, tool arguments, or a complete transcript. Keep the actual `delegation.id` unchanged. Assemble the agent's input from recent role-labeled transcript fragments and verified application state, including the active task and latest correction. A delegation can arrive before a complete sentence appears in the transcript. If the available context does not establish the request, gather more context or ask for clarification before acting.258Use this notification to start your application’s delegation handler. Preserve `delegation.id` so you can attach the result to the same request. Your handler prepares the agent’s input from caller and assistant transcripts plus the task records your application holds; the notification itself contains no request text or tool arguments.

259 

260For example, the appointment agent might receive:

261 

262> Caller: “Actually, Friday instead.”

263>

264> Current request: Find an appointment on Friday in the caller’s time zone.

265>

266> Previous result: Thursday at 2 PM was offered.

267>

268> Confirmation: No Friday slot has been confirmed.

269>

270> Task revision: 2.

271 

272The notification may arrive before the full sentence is transcribed. Keep it until you have enough context, or ask the caller to clarify before taking action.

255 273 

256In a text application, you might pass the user's latest message directly to your agent. With GPT-Live, add an adapter that supplies that context and returns a concise, verified result:274In a text application, you might pass the user's latest message directly to your agent. With GPT-Live, add an adapter that supplies that context and returns a concise, verified result.

275 

276Before calling the adapter, record the delegation ID so only one handler starts work for it. If the request is unclear, call the adapter again when the context is ready.

277 

278Your backend remains responsible for authorization, confirmation, operation IDs, and retries. Check the task’s current revision before changing a booking. The adapter’s later revision check only prevents an outdated result from being announced; it cannot undo a booking already made.

257 279 

258Connect a client delegation to your agent280Connect a client delegation to your agent

259 281 


329```351```

330 352 

331 353 

332The adapter uses application callbacks to read context, run your agent, and check the current task revision; these aren't SDK methods. The context callback returns a ready snapshot containing recent conversation and the current task, or no snapshot when the request remains unclear. The agent callback invokes your existing agent and returns a verified summary of at most 500 tokens. In JavaScript, the application-provided `send` callback sends the JSON event on your Live connection. In Python, the adapter sends the update through the SDK `connection` directly.354Implement the context and agent callbacks in your application. The context callback returns recent conversation and the current task, or no value while the request is still unclear. The agent callback runs your existing backend and returns a verified summary of at most 500 tokens. In JavaScript, the application-provided `send` callback sends the JSON event on your Live connection. In Python, the adapter sends the update through the SDK `connection` directly.

333 

334If context is not ready, retain the notification and invoke the adapter again after resolving the request. Before invoking this adapter, claim the delegation in your application so duplicate delivery cannot start the same operation twice. Keep authorization, confirmation, operation IDs, and retry decisions in your backend. The revision check prevents this adapter from announcing an outdated result; the backend must also check the current revision before a side effect such as booking.

335 355 

336For the appointment assistant, the context should establish the requested date and time zone, previously offered slots, any confirmed slot, and the latest correction. An availability result should say that a slot is available and that no booking has been made. Only return a booking confirmation after the booking succeeds. See [Client delegation](https://developers.openai.com/api/docs/guides/live-delegation#receive-a-client-delegation) for the full setup and result flow.356For the appointment assistant, the context should establish the requested date and time zone, previously offered slots, any confirmed slot, and the latest correction. An availability result should say that a slot is available and that no booking has been made. Only return a booking confirmation after the booking succeeds. See [Client delegation](https://developers.openai.com/api/docs/guides/live-delegation#receive-a-client-delegation) for the full setup and result flow.

337 357 


343- Use `session.commentary.append` for a verified result the user should hear.363- Use `session.commentary.append` for a verified result the user should hear.

344- Use `session.instructions.append` for application-authored behavioral guidance.364- Use `session.instructions.append` for application-authored behavioral guidance.

345 365 

346All three take plain-string `content` of at most 500 tokens and require `delegation_id`. Use the original client delegation ID for related work or `null` for general session context. Match append acknowledgments through `client_event_id`. Acceptance does not establish speech or playback. See [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update).366All three take plain-string `content` of at most 500 tokens and require `delegation_id`. Use the original client delegation ID for related work or `null` for general session context. Match each acknowledgment to the command you sent using `client_event_id`. This confirms that the update was accepted. Use assistant transcript events to observe generated speech and your player’s state to track playback. See [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update).

347 367 

348When the user says “Actually, Friday instead,” update the active task and its revision, invalidate any Thursday confirmation, and direct the existing agent to the corrected request. Decide whether to cancel, change, or let the pending lookup finish. Discard an outdated result before returning it to GPT-Live. An interruption in speech does not cancel a backend operation, and a cancellation request does not prove that an action was cancelled.368When the user says “Actually, Friday instead,” save Friday as the current request, advance its version number, and clear any confirmation for Thursday. Send that correction to your existing agent. Decide whether to request cancellation of the Thursday lookup, change it, or let it finish and discard its result.

369 

370Track the status of the lookup before reporting it as cancelled. Handle that backend decision even if the user’s interruption has already stopped the assistant’s speech.

349 371 

350Backend work may outlive the voice session. Persist its status in your application. In a later voice interaction, start a new session with the relevant saved context; see [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).372Backend work may outlive the voice session. Persist its status in your application. In a later voice interaction, start a new session with the relevant saved context; see [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).

351 373 

374### Keep typed input connected to your agent

375 

376Keep typed input connected to your existing backend. Treat a typed correction as an update to the same task, and send relevant verified context to the voice session. See [Accept typed input](https://developers.openai.com/api/docs/guides/live-delegation#accept-typed-input) and [Keep updates accurate and useful](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful).

377 

352### Adapt text and speech safeguards378### Adapt text and speech safeguards

353 379 

354A text agent can finish and validate a reply before displaying it. A chained pipeline may validate the complete reply before sending it to text-to-speech. GPT-Live can speak while backend work is still running, so withholding a tool result or backend continuation does not hold all speech.380A text agent or chained pipeline can validate a complete reply before displaying or speaking it. With GPT-Live, conversation and backend work run at the same time. If every spoken response must pass a check before the user hears it, put that check in the audio playback path your application controls. Holding a backend result alone will not pause all speech.

355 381 

356Follow [Adapt your guardrails](#adapt-your-guardrails) to retain your checks and account for continuous speech.382Follow [Adapt your guardrails](#adapt-your-guardrails) to retain your checks and account for continuous speech.

357 383 

358Keep typed input connected to your existing backend. Treat a typed correction as an update to the same task, and send relevant verified context to the voice session. See [Accept typed input](https://developers.openai.com/api/docs/guides/live-delegation#accept-typed-input) and [Keep updates accurate and useful](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful).

359 

360 384 

361 385 

362## Adapt your guardrails386## Adapt your guardrails

363 387 

364Keep the input and output safeguards from your existing application when migrating from either architecture. GPT-Live can continue speaking while backend work and policy checks run, so apply checks to both the conversation and the actions your backend takes.388Keep the input and output safeguards from your existing application when migrating from either architecture. GPT-Live can continue speaking while backend work and policy checks run, so apply checks to both the conversation and the actions your backend takes.

365 389 

366Use a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) when your server needs independent access to a browser-owned session. Your server can receive transcripts and send corrective instructions while audio stays on WebRTC. If it already owns the primary WebSocket, use that event stream; Responses delegation does not require an additional sideband.390If your server needs to monitor or control a browser’s WebRTC session, attach a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) to receive transcripts and send corrective instructions. Audio continues over WebRTC. If your server already streams audio through the primary WebSocket, use that connection’s event stream for these checks. Choosing Responses delegation does not by itself require a sideband.

367 391 

3681. Monitor user and assistant transcript events and run your checks alongside the conversation.3921. Monitor user and assistant transcript events and run your checks alongside the conversation.

3692. Block affected tools and external actions in application code. Cancel related application-owned work where supported, and prevent late results from continuing a blocked request.3932. Block affected tools and external actions in application code. Cancel related application-owned work where supported, and prevent late results from continuing a blocked request.


371 395 

372For example, if a caller asks the appointment assistant to change another person's booking without permission, block the booking operation before it runs. Then instruct the assistant to explain that it cannot make the change. Verify both the unchanged booking record and the spoken response; the refusal alone does not enforce authorization.396For example, if a caller asks the appointment assistant to change another person's booking without permission, block the booking operation before it runs. Then instruct the assistant to explain that it cannot make the change. Verify both the unchanged booking record and the spoken response; the refusal alone does not enforce authorization.

373 397 

374A corrective instruction cannot retract audio already heard. If checks must finish before playback, add buffering and approval to the audio path your application controls and account for the added latency. Follow [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for the complete flow, a corrective instruction example, and playback controls. For required opening wording, see [Deliver a disclosure](https://developers.openai.com/api/docs/guides/live-conversations#deliver-a-disclosure).398A corrective instruction cannot retract audio already heard. If your application needs to check assistant speech before playback, follow [Check speech before playback](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#check-speech-before-playback) for buffering, approval, interruption, and recovery handling. See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for action controls and a corrective instruction example. For required opening wording, see [Deliver a disclosure](https://developers.openai.com/api/docs/guides/live-conversations#deliver-a-disclosure).

375 399 

376## Validate the migration400## Validate the migration

377 401 


380- **Actions and spoken confirmations:** Check availability, ask for confirmation, and book only the confirmed slot. Verify the backend outcome, spoken answer, and client playback separately.404- **Actions and spoken confirmations:** Check availability, ask for confirmation, and book only the confirmed slot. Verify the backend outcome, spoken answer, and client playback separately.

381- **Corrections and duplicate prevention:** Change Thursday to Friday during a pending request. Discard outdated results and ensure retries cannot create a second booking.405- **Corrections and duplicate prevention:** Change Thursday to Friday during a pending request. Discard outdated results and ensure retries cannot create a second booking.

382- **Permissions:** Try an unauthorized action and a booking without confirmation. Check that application policy blocks execution.406- **Permissions:** Try an unauthorized action and a booking without confirmation. Check that application policy blocks execution.

383- **Guardrail interventions:** Trigger a check during speech and during tool execution. Verify corrective speech, blocked actions, late-result handling, and playback recovery. Include slow checks and false positives.407- **Guardrail interventions:** Trigger checks during speech and tool execution. Verify that the assistant receives the correction, affected actions stay blocked even if a tool result arrives late, and playback resumes as intended. Check whether a running operation actually stopped. Include slow checks and false positives.

384- **Interruptions:** Speak while the assistant is talking or working. Verify the conversation, audio playback, and backend task state independently.408- **Interruptions:** Speak while the assistant is talking or working. Verify the conversation, audio playback, and backend task state independently.

385- **Failures and reconnects:** Exercise tool errors, lost results, and disconnects. Reconcile uncertain outcomes before retrying, and restore relevant saved context in a new session.409- **Failures and reconnects:** Test tool errors, lost results, and disconnects. If a booking request loses its response, check whether the booking succeeded before retrying. Start the next voice session with the saved task context and verify that it continues from the established outcome.

386 410 

387Use [Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency) to tune the migrated backend. Compare useful spoken response time and task success with the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation), and use [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost) to compare usage and cost.411Use [Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency) to tune the migrated backend. Compare useful spoken response time and task success with the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation), and use [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost) to compare usage and cost.

Details

10| --------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |10| --------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |

11| [LiveKit](https://docs.livekit.io/agents/models/realtime/plugins/gpt-live) | Build GPT-Live voice agents with LiveKit’s OpenAI plugin. |11| [LiveKit](https://docs.livekit.io/agents/models/realtime/plugins/gpt-live) | Build GPT-Live voice agents with LiveKit’s OpenAI plugin. |

12| [Twilio](https://www.twilio.com/en-us/blog/developers/twilio-openai-gpt-live-1-api-resources) | Connect incoming and outgoing phone calls to GPT-Live with Twilio Agent Connect. |12| [Twilio](https://www.twilio.com/en-us/blog/developers/twilio-openai-gpt-live-1-api-resources) | Connect incoming and outgoing phone calls to GPT-Live with Twilio Agent Connect. |

13| [Telnyx](https://telnyx.com/resources/outbound-ai-calls-python-openai-live) | Build outbound calling experiences with GPT-Live and the Telnyx Voice API. |13| [Telnyx](https://developers.telnyx.com/docs/voice/sip-trunking/gpt-live-configuration-guide) | Build outbound calling experiences with GPT-Live and the Telnyx Voice API. |

14| [Daily/Pipecat](https://docs.pipecat.ai/api-reference/server/services/s2s/openai-live) | Add GPT-Live to your application with Pipecat’s OpenAI Live service. |14| [Daily/Pipecat](https://docs.pipecat.ai/api-reference/server/services/s2s/openai-live) | Add GPT-Live to your application with Pipecat’s OpenAI Live service. |

15 15 

16## Integration checklist16## Integration checklist

Details

4 4 

5`gpt-live-1` is a voice model for natural, continuous conversation. It can listen and speak at the same time, respond to interruptions, and keep the conversation moving while a backend agent handles reasoning, tools, and longer tasks.5`gpt-live-1` is a voice model for natural, continuous conversation. It can listen and speak at the same time, respond to interruptions, and keep the conversation moving while a backend agent handles reasoning, tools, and longer tasks.

6 6 

7Give GPT-Live a goal and room to conduct the conversation. The live prompt need not prescribe every question or acknowledgment. Define the assistant’s role, conversational style, and when to involve the backend. Give GPT-Live flexibility in its phrasing, acknowledgments, and pacing.7Use `session.instructions` for the assistant’s role, speaking style, and when to ask the backend for help. Give the backend model or agent the procedures and tools for tasks such as looking up an order or changing a booking.

8 8 

9When migrating from Realtime, start with a simpler prompt. Test which rules for exact wording, fixed response sequences, or turn-taking your product still needs. Revise existing instructions and remove conflicts as you iterate.9Describe the conversational behavior you want, and let GPT-Live choose the wording for ordinary replies. When migrating from Realtime, keep the rules your product needs for wording, interruptions, and the order of actions. Test the simpler prompt on representative conversations as you revise it.

10 10 

11Keep detailed procedures in the backend prompt and enforce permissions and tool execution checks in your application.11Your application checks permissions and required confirmations before executing an action.

12 12 

13## Recommended prompt structure13## Recommended prompt structure

14 14 

15The live model has a small context window. Use the template below as your `session.instructions` value and add only the optional controls your application needs.15Start with this template and add instructions as needed. See [session configuration](https://developers.openai.com/api/docs/guides/live-conversations#configuration-fields) for field limits.

16 16 

17GPT-Live delegates reasoning and tool use to your backend while it handles the conversation. Configure backend prompts and tools in [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation).17Keep the template’s `Backchannel policy`, `Interruption policy`, and `Delegation policy` headings, and customize the text beneath them. Backchannels are brief listening sounds, such as “mm-hmm,” that the assistant can make while the caller continues speaking.

18 

19Keep the policy labels. Customize the personality, backchannel behavior, backend capabilities, and delegation conditions for your product.

20 18 

21```text19```text

22You are [name], a calm, friendly voice assistant for [service].20You are [name], a calm, friendly voice assistant for [service].


43Do not guess the result while waiting.41Do not guess the result while waiting.

44```42```

45 43 

46List only capabilities your backend actually has. These describe what it can help with; they are not instructions for the live model to call a tool.44List the capabilities your backend supports. GPT-Live uses this list to decide which requests to hand off. Configure the backend’s actual tools in [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation).

47 45 

48## Personality46## Personality

49 47 

50Give the assistant a clear role, tone, and pace. Also describe how it should respond when someone is frustrated or unsure. A few short sentences, like the opening of the starter prompt, are enough.48Describe the assistant’s role, tone, and speaking pace in a few sentences. Include how it should respond when a caller is frustrated or unsure. For example: “Explain one step at a time. If the caller sounds confused, ask which part they want to go over.”

51 

52The live prompt controls speaking behavior, including tone, pace, backchannels, and interruptions. Keep long business procedures in the backend prompt.

53 49 

54## Backchannels50## Backchannels

55 51 

56A backchannel is a short listening sound, such as “mm-hmm.” Start with moderate backchannels so the assistant shows it is listening without taking over the conversation.52Start with the template’s backchannel policy, then listen for whether the assistant’s brief acknowledgments help the conversation or interrupt the caller. “Moderate” is a prompting instruction, not a numerical frequency setting.

57 53 

58You can modify this line from the starter prompt:54You can modify this line from the starter prompt:

59 55 


61Backchannel policy: Use moderate backchannels. Acknowledge naturally without competing with the main response.57Backchannel policy: Use moderate backchannels. Acknowledge naturally without competing with the main response.

62```58```

63 59 

64Do not add a blanket “never speak while the user is speaking” rule alongside it. That can also suppress helpful listening sounds. Change the policy only if your product needs different behavior, then listen to real conversations to check the result.60If you want backchannels, allow brief listening sounds during interruptions. A rule that forbids all overlapping speech can suppress them.

65 61 

66## Interruptions62## Interruptions

67 63 

68When the user interrupts, the assistant should stop its answer and listen. A brief listening sound is different from taking over the user's turn.64When the user interrupts, the assistant should stop its answer and listen. A brief listening sound is different from taking over the user's turn.

69 65 

70Stopping speech does not automatically stop backend work. “Stop talking” and “Cancel my booking” mean different things. If the user changes or cancels a request, the backend must handle that change and confirm what happened. See [task state and interruptions](https://developers.openai.com/api/docs/guides/live-delegation).66Handle changes to a task separately from interruptions to speech. “Stop talking” asks the assistant to yield; “Cancel my booking” asks the backend to take an action. Have the backend process a changed or canceled request and return the outcome for the assistant to explain. See [task state and interruptions](https://developers.openai.com/api/docs/guides/live-delegation).

71 67 

72## Delegation68## Delegation

73 69 

74Organize the `Delegation policy` section of your prompt under three labels: `Backend tools`, `Delegate to the backend when`, and `Do not delegate to the backend when`. Describe the backend's capabilities, then give concrete conditions, such as “the user asks to change a booking,” instead of “delegate when needed.”70In the `Delegation policy` section, list the backend’s capabilities and the requests that should trigger a handoff. Use concrete conditions, such as “the user asks to change a booking.” Keep the template’s three labels: `Backend tools`, `Delegate to the backend when`, and `Do not delegate to the backend when`.

75 71 

76Tell GPT-Live when to delegate and what the backend can help with. Put tool-call instructions and result-handling procedures in the backend prompt.72For a booking assistant, replace the starter template’s entire delegation section with:

77 

78For example, replace the starter prompt's delegation section with a policy like this; do not add a second policy:

79 73 

80```text74```text

81Delegation policy:75Delegation policy:


95Do not guess the result while waiting.89Do not guess the result while waiting.

96```90```

97 91 

98List only capabilities your backend has. Check the policy against a few real user requests: which ones should trigger delegation, and which should not?92Test the policy with requests that need backend work, conversational replies the voice model can handle, and corrections to work already in progress.

93 

94Put the full task procedure in the backend instructions and define tools in the backend’s tool configuration. Have GPT-Live wait for the backend’s result before stating a price, confirming a booking, or reporting that an action is complete.

99 95 

100Keep the full procedure and tool schemas in the backend prompt. The live model only needs the short handoff rules. It must not promise a booking, guess a price, or claim an action has finished before the backend confirms it.96You can prompt GPT-Live to acknowledge a request while delegated work runs. As background work progresses, use [`session.commentary.append`](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful) to provide updates you want GPT-Live to say aloud.

101 97 

102For backend prompts, conversation context, tool results, typed input, and API examples, read [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation). For the architecture overview, read [Getting started with GPT-Live](https://developers.openai.com/api/docs/guides/live).98For backend prompts, conversation context, tool results, typed input, and API examples, read [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation). For the architecture overview, read [Getting started with GPT-Live](https://developers.openai.com/api/docs/guides/live).

103 99 

104## Appendix: Optional controls100## Appendix: Optional controls

105 101 

106**Only add a rule if you need to change a specific behavior.** Most applications should start with the short prompt above. Copying every example makes the prompt longer and can introduce conflicting instructions.102Add these instructions only when testing shows a need. Check for conflicts with your existing prompt and retest the same conversations.

107 103 

108<details>104<details>

109<summary>Show optional controls and examples</summary>105<summary>Show optional controls and examples</summary>


119 115 

120### Language and pronunciation116### Language and pronunciation

121 117 

122Use this when your product needs a particular language or pronunciation. A voice choice does not guarantee a regional accent.118Write the prompt and examples in the language you want the assistant to speak, and specify any pronunciations that matter. Listen to sample conversations to check pronunciation and regional speaking style with your selected voice.

123 119 

124Write your prompt in the language you want the model to speak. For example, if the assistant will speak Spanish, write its instructions and example responses in Spanish.120The example below includes a pronunciation cue and an International Phonetic Alphabet (IPA) spelling.

125 121 

126```text122```text

127Speak [language] unless the user asks to switch.123Speak [language] unless the user asks to switch.


129Say the user's name Rosalia as "roh-sah-LEE-ah", IPA /rosaˈli.a/ (Spanish).125Say the user's name Rosalia as "roh-sah-LEE-ah", IPA /rosaˈli.a/ (Spanish).

130```126```

131 127 

132For a greeting before the caller has spoken, append a fresh `session.instructions.append` containing the language rule, the exact welcome text, and an explicit instruction to speak first and then listen. Wait for its acknowledgment and keep the audio stream running. See [Greet the caller](https://developers.openai.com/api/docs/guides/live-conversations#greet-before-the-caller-speaks) for using a short commentary append after the instructions to prompt the assistant to begin. Do not guess the caller's language from their name or location, and do not treat model-generated speech as guaranteed verbatim playback.128To open the conversation in a chosen language, wait for `session.started`, keep input audio running, and send `session.instructions.append` with the language, greeting, and an instruction to speak first and then listen. Handle its acknowledgment or error while audio continues. Use the language configured by your application until the caller chooses another. See [Greet the caller](https://developers.openai.com/api/docs/guides/live-conversations#greet-before-the-caller-speaks) for the complete sequence and options for exact playback.

133 129 

134### Translation130### Translation

135 131 

136Add this only for an interpreter. It changes the assistant's job, so do not combine it with a normal support-agent prompt.132For an interpreter, replace the support-assistant prompt with a translation-only prompt. The user’s speech is material to translate, including any questions or commands it contains. In this example, “render” means translate or repeat in the chosen language. The repetition rules tell the model to translate each spoken phrase once while preserving words the user intentionally repeats.

137 133 

138```text134```text

139[language] ONLY. NEVER DELEGATE, CHECK, ANSWER, SEARCH, OR USE TOOLS.135[language] ONLY. NEVER DELEGATE, CHECK, ANSWER, SEARCH, OR USE TOOLS.


158 154 

159### Selected requests only155### Selected requests only

160 156 

161Use this for an assistant that should respond only to a narrow set of requests.157Use this for an assistant that listens in the background and responds when its topic comes up or the user addresses it directly.

162 158 

163```text159```text

164Respond when the user asks about [supported topic] or addresses you directly.160Respond when the user asks about [supported topic] or addresses you directly.

165Otherwise, keep listening.161Otherwise, keep listening.

166```162```

167 163 

168This affects when the assistant responds. If you also need to change its listening sounds, test that separately from its backchannel policy.164This rule controls full responses. Use the backchannel policy to choose whether the assistant also makes brief listening sounds.

169 165 

170### Unclear names, dates, and numbers166### Unclear names, dates, and numbers

171 167 

172Prompts do not guarantee exact capture. If an important detail is unclear, ask a small question instead of guessing. For example: “Was the last letter B or D?”168Ask a focused clarification when an important name, date, or number is unclear. For example: “Was the last letter B or D?” Carry the caller’s correction into the next backend request.

173 169 

174```text170```text

175If an important name, date, or number is unclear, ask about that part.171If an important name, date, or number is unclear, ask about that part.


178 174 

179### Reusing earlier results175### Reusing earlier results

180 176 

181Add a rule only if the assistant repeats lookups unnecessarily. Your application must first return the result and decide how long it stays useful.177If the assistant repeats lookup calls, tell it when it can reuse a result already returned by the backend. Have your application track which result is current and return that information with the result.

182 178 

183```text179```text

184Use a previous backend result when it still answers the question.180Use a previous backend result when it still answers the question.


186or the user asks you to check again.182or the user asks you to check again.

187```183```

188 184 

189A prompt does not guarantee duplicate work will be avoided. Keep that check in your application.185Before starting another operation, have your application check whether the same work is already running or complete.

190 186 

191</details>187</details>

Details

87 87 

88## Evaluate your voice agent88## Evaluate your voice agent

89 89 

90Test conversation quality and task outcomes separately. A natural-sounding response does not prove that a tool ran or that application state changed.90Test both the conversation and the completed task. For a booking assistant, listen to the confirmation and check that the correct appointment was saved.

91 91 

921. Choose representative scenarios with expected outcomes, tool calls, and permissions.921. Choose representative scenarios with expected outcomes, tool calls, and permissions.

932. Save the audio, events, tool results, and application state needed to verify each outcome. Distinguish a failed evaluation run from a valid run in which the agent fails the task.932. Save the audio, events, tool results, and application state needed to verify each outcome. Distinguish a failed evaluation run from a valid run in which the agent fails the task.


110 110 

111For a GPT-Live evaluation harness, see the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation).111For a GPT-Live evaluation harness, see the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation).

112 112 

113For a Realtime evaluation harness and worked examples, use the [Realtime evaluation guide in the OpenAI Cookbook](https://developers.openai.com/cookbook/examples/realtime_eval_guide). The Cookbook owns the runnable evaluation recipes; this page provides the shared testing checklist.113For a Realtime evaluation harness and worked examples, use the [Realtime evaluation guide in the OpenAI Cookbook](https://developers.openai.com/cookbook/examples/realtime_eval_guide).

114 114 

115### Measure latency115### Measure latency

116 116 

117Define an observed start and end event for every latency metric. Time to first117Measure how long callers wait for a useful spoken answer. Track backend time

118audible response, time to delegation, interruption yield, backend completion,118separately to find delays, and compare the median and 95th percentile across

119and verified task completion measure different boundaries. Use one monotonic119similar calls.

120timeline and report the eligible population, median, and tail latency. Do not

121substitute a backend-only timer for end-to-end response time.

122 120 

123Keep the caller, recording, backend model, prompt, transport, audio cadence, and121Keep the caller, recording, backend model, prompt, transport, audio cadence, and

124grader fixed when comparing frontend models.122grader fixed when comparing frontend models.


129direct visibility into its backend requests; Responses delegation exposes nested127direct visibility into its backend requests; Responses delegation exposes nested

130response events and the custom tools your application runs.128response events and the custom tools your application runs.

131 129 

132Use the intervals to locate delays in connection setup, model work, tools,130Use these timings to find delays in connection setup, model work, tools,

133application buffering, and playback. Measure the first useful spoken answer131buffering, or playback. Measure acknowledgments such as “I'm checking” separately

134separately from an acknowledgment such as “I'm checking.” An earlier132from the answer the caller needs.

135acknowledgment does not show that the requested result arrived sooner.

136 133 

137Change one factor at a time and repeat the same scenarios. Compare median and134Change one factor at a time and repeat the same scenarios. Check whether faster

138tail time to useful spoken responses alongside task success, tool correctness,135responses also affect task success, tool correctness, or interruptions. See

139and interruptions. See [Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency)136[Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency).

140for implementation guidance.

141 137 

142## Voice agents still use the same core agent building blocks138## Voice agents still use the same core agent building blocks

143 139 

Details

109session is open or closed. Save the task state and conversation context before109session is open or closed. Save the task state and conversation context before

110[closing the voice session](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close).110[closing the voice session](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close).

111 111 

112For the inactivity timeout, restart trigger, and context restoration workflow, see [Close idle sessions and resume](https://developers.openai.com/api/docs/guides/live-conversations#close-idle-sessions-and-resume).

113 

112For an ambient agent, close the voice session while the backend handles a114For an ambient agent, close the voice session while the backend handles a

113long-running task, such as coding in goal mode. Offer a button labeled115long-running task, such as coding in goal mode. Offer a button labeled

114**Resume conversation** to start a new voice session when the user returns, or use a backend116**Resume conversation** to start a new voice session when the user returns, or use a backend

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.5Choose your API to see its connection steps and session events.

6 6 

7 7 

8 8 


16 16 

17For browser applications, use the [WebRTC data channel](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for captions and local UI updates. Use a sideband when transcript processing runs on your server, such as guardrail checks, sentiment analysis, or speculative tool calls. Your server can receive events and steer the same session directly while browser audio stays on WebRTC. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples.17For browser applications, use the [WebRTC data channel](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for captions and local UI updates. Use a sideband when transcript processing runs on your server, such as guardrail checks, sentiment analysis, or speculative tool calls. Your server can receive events and steer the same session directly while browser audio stays on WebRTC. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples.

18 18 

19If your backend already owns the primary [WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live), it already receives the session's events and can send commands.19If your backend streams audio over the primary [WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live), use that connection to receive events and send commands.

20 20 

21[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation) also works without a sideband. The browser can forward function-call events from its data channel to an authenticated backend for execution. OpenAI-hosted tools run through the delegated backend without an application tool executor.21[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation) also works without a sideband. The browser can forward function-call events from its data channel to an authenticated backend for execution. OpenAI-hosted tools run through the delegated backend without an application tool executor.

22 22 


31 31 

323. Receive events and send commands on the attached socket. The session is already running; do not send `session.start` again.323. Receive events and send commands on the attached socket. The session is already running; do not send `session.start` again.

33 33 

34Treat the session ID as an opaque value. Preserve its prefix and use it only for the session to which your application has authorized access. Read the ID from the Live JSON response, rather than a Realtime `Location` header or `call_id` URL parameter.34Use the session ID unchanged, including its prefix, and verify that your application has authorized access to that session.

35 35 

36### Observe events and send commands36### Observe events and send commands

37 37 


46 46 

47Commands follow the same validation and delegation rules as on the primary connection. For context appends, use `delegation_id: null` for general session context; a non-null ID must identify an existing client delegation. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for configuration, function execution, and the context append examples.47Commands follow the same validation and delegation rules as on the primary connection. For context appends, use `delegation_id: null` for general session context; a non-null ID must identify an existing client delegation. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for configuration, function execution, and the context append examples.

48 48 

49For browser sessions, keep microphone input and speaker output on the negotiated WebRTC media track. Use the sideband for conversation events and control. A transcript event or command acknowledgment does not prove that audio has played or that the user has heard it.49For browser sessions, keep microphone input and speaker output on the negotiated WebRTC media track. Use the sideband for conversation events and control, and track playback in your audio player.

50 50 

51### Receive reflected audio51### Receive reflected audio

52 52 


59 59 

60Both payloads are base64-encoded raw mono PCM16LE at 24 kHz, regardless of the primary transport's audio format. Neither event has an `event_id`. Reflected input contains received audio before input muting; it does not confirm that the model consumed those samples. Reflected output ranges can have gaps for dropped frames and do not indicate when the caller heard the audio.60Both payloads are base64-encoded raw mono PCM16LE at 24 kHz, regardless of the primary transport's audio format. Neither event has an `event_id`. Reflected input contains received audio before input muting; it does not confirm that the model consumed those samples. Reflected output ranges can have gaps for dropped frames and do not indicate when the caller heard the audio.

61 61 

62These are server events, not permission to send audio through the sideband. Send microphone audio through the primary transport; do not send `session.input_audio.append` on the attached socket.62Send microphone audio only through the primary connection. Use the sideband to receive reflected audio.

63 63 

64### Assign one owner for each action64### Assign one owner for each action

65 65 

66Choose whether the browser or backend handles each action. If both connections receive a function-call event, execute the function once. Apply the same ownership rule to context updates and requests to continue backend work.66Choose whether the browser or backend handles each action. If both connections receive a function-call event, execute the function once. Apply the same ownership rule to context updates and requests to continue backend work.

67 67 

68Store transcripts and tool state in your application. Attach early if the backend needs to observe the conversation from the start, and retain any history collected before attachment. Do not rely on attachment to reconstruct earlier transcripts or tool results.68Attach early if the backend needs to observe the conversation from the start. Store transcripts and tool state in your application, including any history collected before attachment.

69 69 

70A sideband does not itself make session events private from the browser. Keep sensitive tool credentials and authorization decisions in your backend, and return only the context needed for the conversation.70The browser can still receive session events when a sideband is attached. Keep sensitive tool credentials and authorization decisions in your backend, and return only the context needed for the conversation.

71 71 

72 72 

73 73 


75 75 

76## Apply conversation guardrails76## Apply conversation guardrails

77 77 

78Use your server's connection to monitor the conversation, check requests against your application's policies, and intervene when a check triggers. A sideband gives your server access to session events and commands; your application runs the checks and enforces their results. The same workflow applies when your server already owns the primary WebSocket connection.78Use your server's primary WebSocket or sideband to monitor the conversation and check requests against your application's policies. Your application runs the checks, blocks affected actions, and sends corrective instructions when a check triggers.

79 79 

80### Run checks alongside the conversation80### Run checks alongside the conversation

81 81 

82Guardrails are one use of [processing transcript fragments as they arrive](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments). The same stream can start a speculative lookup or update the UI alongside these checks.82Guardrails are one use of [processing transcript fragments as they arrive](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments). The same stream can start a speculative lookup or update the UI alongside these checks.

83 83 

841. **Monitor transcripts.** Accumulate `session.input_transcript.delta` fragments to check user requests for jailbreak attempts, sensitive information, or policy violations. Use `session.output_transcript.delta` to check assistant speech for unsupported claims or responses outside your application's scope. Keep each check associated with the transcript and application request it evaluated.841. **Monitor transcripts.** Accumulate `session.input_transcript.delta` fragments to check user requests for jailbreak attempts, sensitive information, or policy violations. Use `session.output_transcript.delta` to check assistant speech for unsupported claims or responses outside your application's scope. Keep each check associated with the transcript and application request it evaluated.

852. **Run checks concurrently.** A fast, lightweight model can evaluate requests while the conversation continues. Return a small structured result, such as `{"triggered": true}`, that your application can act on. Keep actions that require approval blocked until their checks pass; a timeout or failed check is not approval.852. **Run checks concurrently.** A fast, lightweight model can evaluate requests while the conversation continues. Return a small structured result, such as `{"triggered": true}`, that your application can act on. Run actions that require approval only after their checks pass. Keep them blocked if a check fails or times out.

863. **Block affected actions.** When a check triggers, mark the request as blocked in application state. Check that state before executing a tool or committing a change, including work already queued. A spoken refusal does not prevent a tool from running.863. **Block affected actions.** When a check triggers, mark the request as blocked in application state. Check that state before executing a tool or committing a change, including work already queued.

874. **Stop related work.** Cancel application-owned jobs where your backend supports cancellation, and discard late results from blocked or superseded requests. With Responses delegation, stop executing affected custom functions and do not send `response.create` to continue blocked work. This does not cancel an already-running hosted response or stop frontend speech.874. **Stop related work.** Cancel application-owned jobs where your backend supports cancellation, and discard late results from blocked or superseded requests. With Responses delegation, stop executing affected custom functions and do not send `response.create` to continue blocked work. This does not cancel an already-running hosted response or stop frontend speech.

885. **Record and redirect.** Log the decision with the affected request and delegation IDs, then send a corrective instruction. An event name such as `guardrail.triggered` belongs to your application's telemetry; it is not a GPT-Live API event.885. **Record and redirect.** Log the decision with the affected request and delegation IDs, then send a corrective instruction.

89 89 

90See [Transcript deltas](https://developers.openai.com/api/docs/guides/live-conversations#transcript-deltas) for collecting fragments and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful) for keeping backend results aligned with the current task.90See [Transcript deltas](https://developers.openai.com/api/docs/guides/live-conversations#transcript-deltas) for collecting fragments and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful) for keeping backend results aligned with the current task.

91 91 


126 126 

127Keep the instruction application-authored. Do not copy untrusted user text into it as an instruction. Use `delegation_id: null` for this session-wide correction, and keep `content` within 500 tokens.127Keep the instruction application-authored. Do not copy untrusted user text into it as an instruction. Use `delegation_id: null` for this session-wide correction, and keep `content` within 500 tokens.

128 128 

129Match `session.instructions.appended` to your command through `client_event_id`. The acknowledgment arrives after estimated context injection; it does not prove that the assistant stopped speaking or that queued audio stopped playing. Corrective instructions cannot retract audio the user has already heard.129Match `session.instructions.appended` to your command through `client_event_id`. The acknowledgment arrives after estimated context injection. To stop audio reaching the caller, use the playback controls below.

130 130 

131For disclosures that request specific spoken wording, also use instructions. See [Deliver a disclosure](https://developers.openai.com/api/docs/guides/live-conversations#deliver-a-disclosure) for an example and playback considerations.131For disclosures that request specific spoken wording, also use instructions. See [Deliver a disclosure](https://developers.openai.com/api/docs/guides/live-conversations#deliver-a-disclosure) for an example and playback considerations.

132 132 

133### Control playback when needed133### Control playback when needed

134 134 

135Test corrective instructions and action blocking first. If your application also needs to block model audio, control output at the client or media relay: temporarily mute or drop the output, discard locally queued audio, send the corrective instruction, and resume playback according to your application's recovery policy. Clear stale audio before resuming. A sideband alone does not control the media path, and an instruction acknowledgment is not a signal to resume playback.135If your application needs to block model audio, control playback at the client or media relay. Temporarily mute or drop the output, discard locally queued audio, and send the corrective instruction. Resume playback according to your application's recovery policy after clearing stale audio; an instruction acknowledgment is not a signal to resume. A sideband alone does not control playback, and corrective instructions cannot retract audio the caller has already heard.

136 136 

137`session.input_audio.mute` controls the caller's microphone input. It does not mute model output or cancel delegated work.137`session.input_audio.mute` controls the caller's microphone input. It does not mute model output or cancel delegated work.

138 138 

139GPT-Live streams transcript fragments while speaking. If a check must finish before the user hears the audio, your application needs to buffer and approve audio before playback. This adds latency. Suppressed audio can also leave the model's conversation context ahead of what the user heard, so test how the conversation resumes.139### Check speech before playback

140 

141For most applications, [monitor user and assistant transcripts](#run-checks-alongside-the-conversation) while the conversation continues. When a guardrail triggers, your application can block affected actions or send corrective instructions.

142 

143If your application needs to check assistant speech before playback, buffer the audio in your player or media bridge before sending it to the caller. Keep receiving audio and transcript events while the check runs, and read transcripts independently of playback.

144 

1451. **Collect audio and its transcript.** Implement output voice activity detection (VAD) in your application, or use a noise gate provided by your media framework, to identify candidate speech segments. GPT-Live does not provide an output VAD or noise gate for this workflow. Wait for the transcript needed to check each segment.

1462. **Release approved audio.** When a segment passes the check, enqueue its original buffered audio for playback. If the check fails, the transcript is missing, or the check times out, discard the segment and use an application-defined safe fallback.

1473. **Handle interruptions.** Associate the audio, transcript, check result, and playback state with an application-generated ID. When an interruption cancels that speech, clear its buffered and queued audio and ignore any later approval for it.

148 

149A pause can mark a candidate segment while the model is still composing an answer. If your policy requires checking a complete answer, define how your application establishes completion; voice activity detection alone cannot establish that the whole answer is finished.

150 

151Buffering adds latency. Discarded speech remains in the model’s conversation context, so test how the conversation continues after withheld audio or a fallback.

140 152 

141### Test the intervention153### Test the intervention

142 154 


144 156 

145## Finish cleanly157## Finish cleanly

146 158 

147Keep receiving events while the backend owns tool execution or final usage collection. Register the `session.closed` handler before sending `session.close`, and keep the WebRTC connection, data channel, and sideband open while pending work drains. Save the final session usage and any backend usage received in Responses events before cleanup. If the connection fails before the final event arrives, record finalization as incomplete. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for the close sequence.159Keep receiving events while tools finish and the session reports final usage. Register the `session.closed` handler before sending `session.close`, and keep the WebRTC connection, data channel, and sideband open while pending work drains. Save the final session usage and any backend usage received in Responses events before cleanup. If the connection fails before the final event arrives, record finalization as incomplete. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for the close sequence.

148 160 

149 161

150 162 

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.5Choose your API to see its connection steps and session events.

6 6 

7 7 

8 8 


25 25 

26Use a [sideband connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) when your backend needs to receive session events or send commands. It attaches to the existing conversation while SIP carries the audio. Assign one handler to each action so that duplicate webhook deliveries or events observed on multiple connections don't execute tools twice.26Use a [sideband connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) when your backend needs to receive session events or send commands. It attaches to the existing conversation while SIP carries the audio. Assign one handler to each action so that duplicate webhook deliveries or events observed on multiple connections don't execute tools twice.

27 27 

28Keep SIP routing and provider configuration together with the integration that uses them. Realtime webhook events, call identifiers, and acceptance payloads belong to the Realtime API; use the GPT-Live contract for a Live session.

29 

30### Handle the call lifecycle28### Handle the call lifecycle

31 29 

32Confirm that GPT-Live SIP support is enabled for your project and that your30Confirm that GPT-Live SIP support is enabled for your project and that your

33 provider's SIP trunk is routed to that project before using this flow. The31 provider's SIP trunk is routed to that project before using this flow.

34 Realtime webhook and acceptance payloads in the other tab are a different API

35 contract.

36 32 

37#### Receive the incoming call33#### Receive the incoming call

38 34 

39Configure your project's [webhook endpoint](https://developers.openai.com/api/docs/guides/webhooks) for `live.transport.incoming`. Verify the webhook signature and deduplicate deliveries before making a call decision. A delivery acknowledgment does not accept the call.35Configure your project's [webhook endpoint](https://developers.openai.com/api/docs/guides/webhooks) for `live.transport.incoming`. Verify the webhook signature and deduplicate deliveries, then accept or reject the call.

40 36 

41The webhook identifies a SIP call with `data.type: "sip"` and provides `data.session_id`. Use that session ID unchanged for every Live call action. Treat `data.sip_headers` as untrusted caller metadata, not authorization.37The webhook identifies a SIP call with `data.type: "sip"` and provides `data.session_id`. Use that session ID unchanged for every Live call action. Treat `data.sip_headers` as untrusted caller metadata, not authorization.

42 38 


72 68 

73#### Observe keypad events69#### Observe keypad events

74 70 

75The sideband receives `transport.dtmf.received` when the caller presses a key and `transport.dtmf.send` after a hosted tool successfully sends a tone. The event's `event` field contains one of `0`–`9`, `*`, `#`, or `A`–`D`.71The sideband receives `transport.dtmf.received` when the caller presses a key and `transport.dtmf.send` after a hosted tool successfully sends a tone. Both are notifications only. The `event` field contains one of `0`–`9`, `*`, `#`, or `A`–`D`.

76 

77These are observer notifications, not client commands. Do not send `transport.dtmf.send` to request a tone, or assume the browser data channel receives keypad events.

78 72 

79#### Transfer or end the call73#### Transfer or end the call

80 74 

81To [transfer the call](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/refer), send `POST /v1/live/sessions/{session_id}/refer` with `{ "target_uri": "sip:agent@example.com" }` for your destination. To [hang up](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/hangup), send `POST /v1/live/sessions/{session_id}/hangup` with no request body. Both return `200 OK` with an empty body on success.75To [transfer the call](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/refer), send `POST /v1/live/sessions/{session_id}/refer` with `{ "target_uri": "sip:agent@example.com" }` for your destination. To [hang up](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/hangup), send `POST /v1/live/sessions/{session_id}/hangup` with no request body. Both return `200 OK` with an empty body on success.

82 76 

83Keep your sideband open for final events and usage before releasing application resources. A successful hangup request or an unexpected disconnect is not a substitute for `session.closed`. See [Usage and graceful close](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for finalization and close reasons.77Keep the sideband open until `session.closed` supplies final usage, then release application resources. If the connection drops first, record finalization as incomplete. See [Usage and graceful close](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for finalization and close reasons.

84 78 

85This flow accepts inbound calls. Creating an outbound SIP call through `POST /v1/live/sessions` is not supported; use the relevant [partner integration](https://developers.openai.com/api/docs/guides/live-partner-integrations) for provider-owned outbound calling.79This flow accepts inbound calls. Creating an outbound SIP call through `POST /v1/live/sessions` is not supported; use the relevant [partner integration](https://developers.openai.com/api/docs/guides/live-partner-integrations) for provider-owned outbound calling.

86 80 


88 82 

89Use the [GPT-Live WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) when your application receives an audio stream from a phone provider or an agent framework. The application authenticates both connections, translates their event envelopes, and relays audio in both directions.83Use the [GPT-Live WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) when your application receives an audio stream from a phone provider or an agent framework. The application authenticates both connections, translates their event envelopes, and relays audio in both directions.

90 84 

91GPT-Live supports raw G.711 μ-law and A-law audio at 8 kHz over WebSocket. When the provider stream uses the same codec, sample rate, and channel count, your application can forward the raw audio bytes without converting them to PCM. Preserve audio order and use the message format required by each connection. Matching audio formats don't make the two event protocols interchangeable.85GPT-Live supports raw G.711 μ-law and A-law audio at 8 kHz over WebSocket. When the provider stream uses the same codec, sample rate, and channel count, your application can forward the raw audio bytes without converting them to PCM. Preserve audio order and wrap the audio bytes in the message format required by each connection.

92 86 

93The bridge also owns any audio it queues for playback. Include provider buffering, interruptions, and ending the call in your application design. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for the Live session lifecycle and [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration) for changes to turn-taking and playback control.87Have your bridge manage queued audio, interruptions, and call termination. Account for audio buffered by the provider when handling playback. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for the Live session lifecycle and [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration) for changes to turn-taking and playback control.

94 88 

95Keep the provider's call or room identifier alongside the OpenAI session ID so you can trace a conversation across both systems.89Keep the provider's call or room identifier alongside the OpenAI session ID so you can trace a conversation across both systems.

96 90 

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.5Choose your API to see its connection steps and session events.

6 6 

7 7 

8 8 

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.5Choose your API to see its connection steps and session events.

6 6 

7 7 

8 8 


10 10 

11Use a primary WebSocket when your server captures audio or relays an audio stream for a client. It carries audio and JSON events in both directions. Keep the project API key on that trusted server. For browser and mobile applications, start with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live).11Use a primary WebSocket when your server captures audio or relays an audio stream for a client. It carries audio and JSON events in both directions. Keep the project API key on that trusted server. For browser and mobile applications, start with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live).

12 12 

13This guide covers the primary audio connection. A [sideband connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) lets a server observe and control an existing Live session. A [Responses WebSocket](https://developers.openai.com/api/docs/guides/websocket-mode) connects your backend to the Responses API for reasoning and tools. Neither replaces the primary audio connection.13This guide covers streaming audio to GPT-Live. To monitor or control an existing session, see [Server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live). To connect a reasoning and tool backend to the Responses API, see [Responses WebSocket mode](https://developers.openai.com/api/docs/guides/websocket-mode).

14 14 

15### Authenticate and start the session15### Authenticate and start the session

16 16 


263- **Receive backend events:** when using Responses delegation, process the nested `event` in each `response.event` envelope.263- **Receive backend events:** when using Responses delegation, process the nested `event` in each `response.event` envelope.

264- **Handle errors:** handle rejected commands and session errors from `error` events. Use `error.client_event_id`, when present, to identify the command.264- **Handle errors:** handle rejected commands and session errors from `error` events. Use `error.client_event_id`, when present, to identify the command.

265 265 

266Output audio events have no timing fields, and GPT-Live does not emit an output-audio-done event. Track your playback queue to know which received audio has played. Transcript timestamps describe intervals on the session timeline; they do not mark audio playback completion. A backend response completing also does not mean the assistant has finished speaking.266Track playback with your application’s audio queue. GPT-Live’s primary WebSocket sends output audio without timing fields or an output-audio-done event. Use transcript timestamps to organize captions and backend events to track delegated work.

267 267 

268GPT-Live manages when to listen and speak as audio streams. It does not use Realtime's input-buffer commit and `response.create` voice-turn loop. In Live, `response.create` starts or continues delegated backend work. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for that workflow.268GPT-Live manages when to listen and speak as audio streams continuously. Use `response.create` to start or continue delegated backend work. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for that workflow.

269 269 

270### Configure an ongoing session270### Configure an ongoing session

271 271