SpyBara
Go Premium

Documentation 2026-09-09 23:59 UTC to 2026-09-10 18:01 UTC

30 files changed +4,749 −1,470. View all changes and history on the product overview
2026
Tue 29 22:57 Mon 28 22:57 Sat 26 23:59 Fri 25 23:58 Thu 24 23:58 Wed 23 23:58 Tue 22 23:57 Mon 21 23:00 Sat 19 23:00 Fri 18 22:59 Thu 17 10:04 Wed 16 20:58 Tue 15 22:59 Mon 14 22:58 Sun 13 15:02 Fri 11 20:00 Thu 10 18:01 Wed 9 23:59 Sat 5 17:01 Fri 4 23:59 Thu 3 23:00 Wed 2 22:59

guides/audio.md +29 −497

Details

1# Audio and speech1# Audio and voice

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Audio models can understand spoken input, generate spoken output, or do both in the same interaction. This guide explains the vocabulary used across OpenAI's audio docs. When you're ready to choose an implementation path, start with the [Realtime and audio overview](https://developers.openai.com/api/docs/guides/realtime).5For a new conversational voice application, start with **[GPT-Live](https://developers.openai.com/api/docs/guides/live)**. It can listen while speaking and keep the conversation moving while a backend agent reasons, uses tools, or completes a task.

6 6 

7## Audio modalities7Connect your first conversation with the [WebRTC quickstart](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live), then write a short [Live prompt](https://developers.openai.com/api/docs/guides/live-prompting). If you already have a Realtime application or text agent, follow [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration).

8 8 

9An audio application combines one or more of these modalities:9## Choose another audio workflow

10 10 

11| Modality | Meaning | Common use cases |

12| --------------- | -------------------------------------------- | ------------------------------------------------- |

13| Audio input | The model receives sound from a user or app. | Voice agents, transcription, translation. |

14| Audio output | The model or API returns spoken audio. | Voice agents, text to speech, spoken responses. |

15| Text transcript | Speech becomes text. | Captions, call analysis, search, records. |

16| Text prompt | Text controls what the model says or does. | Speech generation, scripted voice flows, prompts. |

17 11 

18## Common speech tasks

19 12 

20**Speech to text** converts speech into text. Use it for captions, notes, transcripts, analytics, search, and accessibility. Transcription can be request-based for files or streaming for live audio. Start with the [Transcription overview](https://developers.openai.com/api/docs/guides/transcription) to choose a workflow and model.

21 13 

22**Text to speech** converts text into spoken audio. Use it for narration, assistants, accessibility, and generated voice responses. Speech generation can stream audio back as the model produces it.

23 14 

24**Speech to speech** lets a model listen, reason, and speak in one low-latency session. Use it for conversational voice agents when the assistant needs to respond, call tools, or maintain session state.

25 15 

26**Speech translation** listens to speech in one language and returns translated speech or transcript output in another language. Use a dedicated realtime translation session when translation should begin continuously as audio arrives.

27 16 

28## Streaming and latency

29 17 

30Streaming means the client and service exchange partial input or output while the interaction is still active. Streaming is useful when users expect immediate feedback, such as live captions, calls, voice agents, and translation.

31 18 

32Lower latency requires a realtime connection, more careful audio handling, and a session model that can emit partial events. Request-based APIs are simpler for file uploads and non-interactive work, but they don't support the same live interaction patterns.

33 19 

34## Request-based APIs and realtime sessions

35 20 

36OpenAI supports two broad audio architectures:

37 21 

38| Architecture | Use when | Examples |

39| --------------------------- | ---------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |

40| Request-based audio APIs | You have a file, a text input, or a bounded request. | [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text), [text to speech](https://developers.openai.com/api/docs/guides/text-to-speech). |

41| Realtime sessions | Audio is live and the app needs low-latency events. | [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents), [translation](https://developers.openai.com/api/docs/guides/realtime-translation), [transcription](https://developers.openai.com/api/docs/guides/realtime-transcription). |

42| Multimodal Chat Completions | You are extending an existing chat flow with audio. | [Audio input or output](#add-audio-to-your-existing-application). |

43 22 

44For build-path guidance, see the [Realtime and audio overview](https://developers.openai.com/api/docs/guides/realtime).

45 23 

46## Add audio to your existing application

47 24 

48Models such as [`gpt-realtime-2.1`](https://developers.openai.com/api/docs/models/gpt-realtime-2.1) and [`gpt-audio-1.5`](https://developers.openai.com/api/docs/models/gpt-audio-1.5) are natively multimodal, meaning they can understand and generate audio and text as input and output.

49 25 

50For live browser speech-to-speech interactions, start with a realtime session in the Agents SDK for JavaScript:

51 26 

52Start a realtime voice session27Use the Realtime API when you need its session and tool model. For transcription, translation, or speech generation without a conversational agent, choose the dedicated API below.

28 

29 

30 

31 

53 32 

54```javascript33| Build | Start here | What you control |

55import { RealtimeAgent, RealtimeSession } from "@openai/agents/realtime";34| ------------------------------------------------------------------ | ---------------------------------------------------------------------------- | ----------------------------------------------------------------- |

35| A speech-to-speech agent using the Realtime session and tool model | [Realtime API](https://developers.openai.com/api/docs/guides/realtime) | Audio turns, session state, tools, and interruptions. |

36| A voice interface for an existing text agent | [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents#build-a-chained-voice-workflow) | Speech-to-text, the text-agent workflow, then text-to-speech. |

37| A transcript of an audio file | [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text) | File uploads, bounded requests, and supported transcript formats. |

38| Live captions without assistant speech | [Live transcription](https://developers.openai.com/api/docs/guides/realtime-transcription) | Streaming audio and incremental transcript events. |

39| Continuous speech translation | [Live translation](https://developers.openai.com/api/docs/guides/realtime-translation) | A dedicated translation session, not a voice-agent turn loop. |

40| Narration or generated speech | [Text to speech](https://developers.openai.com/api/docs/guides/text-to-speech) | Text, voice, and output format. |

41| Audio input or output in an existing chat app | [Audio in Chat Completions](https://developers.openai.com/api/docs/guides/audio-chat-completions) | Bounded multimodal chat requests. |

56 42 

57const agent = new RealtimeAgent({43## Build with voice

58 name: "Assistant",

59 instructions: "You are a helpful voice assistant.",

60});

61 44 

62const session = new RealtimeSession(agent, {45Use [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) to compare architectures. Start with the prompting guide for [GPT-Live](https://developers.openai.com/api/docs/guides/live-prompting) or [Realtime](https://developers.openai.com/api/docs/guides/voice-prompting). Then use the shared guides for [custom voices](https://developers.openai.com/api/docs/guides/custom-voices), [evaluation](https://developers.openai.com/api/docs/guides/voice-agents#evaluate-your-voice-agent), and [cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost). Each guide distinguishes model- or API-specific behavior.

63 model: "gpt-realtime-2.1",46 

64});47## Choose a connection

48 

49For browser audio, start with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc). For server audio pipelines, use [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets). For phone calls, see [Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip). A [server-side control connection](https://developers.openai.com/api/docs/guides/voice-server-controls) lets a trusted backend observe and control a media session.

50 

51Select your API on each connection page. Sharing a transport does not make GPT-Live and Realtime handshakes, credentials, or event formats interchangeable. Check the connection guide for prerequisites and setup instructions.

52 

53## Add audio to your existing application

65 54 

66await session.connect({55The Chat Completions examples now live in [Audio in Chat Completions](https://developers.openai.com/api/docs/guides/audio-chat-completions). For a browser voice-agent starter, use the [GPT-Live WebRTC quickstart](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live).

67 apiKey: "ek_...(ephemeral key from your server)",

68});

69```

70 

71 

72This JavaScript example uses the Agents SDK to connect browser voice agents with WebRTC from the client. For Python voice workflows, use the [Voice agents guide](https://developers.openai.com/api/docs/guides/voice-agents), which covers chained voice pipelines.

73 

74If you already have a text-based LLM application with the [Chat Completions endpoint](https://developers.openai.com/api/reference/resources/chat), you may want to add audio capabilities. For example, if your chat application supports text input, you can add audio input and output: include `audio` in the `modalities` array and use an audio model, like [`gpt-audio-1.5`](https://developers.openai.com/api/docs/models/gpt-audio-1.5).

75 

76The [Responses API](https://developers.openai.com/api/reference/resources/responses) docs currently describe

77 text and image inputs with text outputs. For this audio-chat pattern, use Chat

78 Completions with an audio-capable model.

79 

80 

81 

82Audio output from model

83 

84 Create a human-like audio response to a prompt

85 

86```javascript

87import { writeFileSync } from "node:fs";

88import OpenAI from "openai";

89 

90const openai = new OpenAI();

91 

92// Generate an audio response to the given prompt

93const response = await openai.chat.completions.create({

94 model: "gpt-audio-1.5",

95 modalities: ["text", "audio"],

96 audio: { voice: "alloy", format: "wav" },

97 messages: [

98 {

99 role: "user",

100 content: "Is a golden retriever a good family dog?",

101 },

102 ],

103 store: true,

104});

105 

106// Inspect returned data

107console.log(response.choices[0]);

108 

109// Write audio data to a file

110writeFileSync(

111 "dog.wav",

112 Buffer.from(response.choices[0].message.audio.data, "base64"),

113 { encoding: "utf-8" }

114);

115```

116 

117```python

118import base64

119from openai import OpenAI

120 

121client = OpenAI()

122 

123completion = client.chat.completions.create(

124 model="gpt-audio-1.5",

125 modalities=["text", "audio"],

126 audio={"voice": "alloy", "format": "wav"},

127 messages=[{"role": "user", "content": "Is a golden retriever a good family dog?"}],

128)

129 

130print(completion.choices[0])

131 

132wav_bytes = base64.b64decode(completion.choices[0].message.audio.data)

133with open("dog.wav", "wb") as f:

134 f.write(wav_bytes)

135```

136 

137```go

138package main

139 

140import (

141 "context"

142 "encoding/base64"

143 "fmt"

144 "os"

145 

146 "github.com/openai/openai-go/v3"

147)

148 

149func main() {

150 client := openai.NewClient()

151 response, err := client.Chat.Completions.New(context.Background(), openai.ChatCompletionNewParams{

152 Model: "gpt-audio-1.5",

153 Modalities: []string{"text", "audio"},

154 Audio: openai.ChatCompletionAudioParam{

155 Voice: openai.ChatCompletionAudioParamVoiceUnion{OfString: openai.String("alloy")},

156 Format: openai.ChatCompletionAudioParamFormatWAV,

157 },

158 Messages: []openai.ChatCompletionMessageParamUnion{openai.UserMessage("Is a golden retriever a good family dog?")},

159 })

160 if err != nil {

161 panic(err)

162 }

163 fmt.Println(response.Choices[0])

164 audio, err := base64.StdEncoding.DecodeString(response.Choices[0].Message.Audio.Data)

165 if err != nil {

166 panic(err)

167 }

168 if err := os.WriteFile("dog.wav", audio, 0o600); err != nil {

169 panic(err)

170 }

171}

172```

173 

174```java

175import com.openai.client.OpenAIClient;

176import com.openai.client.okhttp.OpenAIOkHttpClient;

177import com.openai.models.chat.completions.ChatCompletionAudioParam;

178import com.openai.models.chat.completions.ChatCompletionCreateParams;

179import java.io.IOException;

180import java.nio.file.Files;

181import java.nio.file.Path;

182import java.util.Base64;

183 

184ChatCompletionCreateParams params =

185 ChatCompletionCreateParams.builder()

186 .model("gpt-audio-1.5")

187 .addUserMessage("Is a golden retriever a good family dog?")

188 .addModality(ChatCompletionCreateParams.Modality.TEXT)

189 .addModality(ChatCompletionCreateParams.Modality.AUDIO)

190 .audio(

191 ChatCompletionAudioParam.builder()

192 .voice("alloy")

193 .format(ChatCompletionAudioParam.Format.WAV)

194 .build())

195 .store(true)

196 .build();

197 

198var message = client.chat().completions().create(params).choices().get(0).message();

199var audio =

200 message.audio().orElseThrow(() -> new IllegalStateException("No audio output returned"));

201Files.write(Path.of("dog.wav"), Base64.getDecoder().decode(audio.data()));

202message.content().ifPresent(System.out::println);

203```

204 

205```csharp

206using OpenAI.Chat;

207#pragma warning disable OPENAI001

208 

209string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;

210ChatClient client = new("gpt-audio-1.5", key);

211 

212ChatCompletionOptions options = new()

213{

214 ResponseModalities = ChatResponseModalities.Text | ChatResponseModalities.Audio,

215 AudioOptions = new(ChatOutputAudioVoice.Alloy, ChatOutputAudioFormat.Wav),

216 StoredOutputEnabled = true,

217};

218 

219ChatCompletion completion = await client.CompleteChatAsync(

220 [new UserChatMessage("Is a golden retriever a good family dog?")],

221 options

222);

223 

224if (completion.OutputAudio is not ChatOutputAudio audio)

225{

226 throw new InvalidOperationException("No audio output was returned.");

227}

228 

229Console.WriteLine(audio.Transcript);

230await File.WriteAllBytesAsync("dog.wav", audio.AudioBytes.ToArray());

231```

232 

233```ruby

234require "base64"

235require "openai"

236 

237client = OpenAI::Client.new

238completion = client.chat.completions.create(

239 model: "gpt-audio-1.5",

240 messages: [{role: :user, content: "Is a golden retriever a good family dog?"}],

241 modalities: [:text, :audio],

242 audio: {voice: :alloy, format: :wav},

243 store: true

244)

245 

246audio = completion.choices.fetch(0).message.audio or raise "No audio returned"

247File.binwrite("dog.wav", Base64.strict_decode64(audio.data))

248```

249 

250```bash

251curl "https://api.openai.com/v1/chat/completions" \

252 -H "Content-Type: application/json" \

253 -H "Authorization: Bearer $OPENAI_API_KEY" \

254 -d '{

255 "model": "gpt-audio-1.5",

256 "modalities": ["text", "audio"],

257 "audio": { "voice": "alloy", "format": "wav" },

258 "messages": [

259 {

260 "role": "user",

261 "content": "Is a golden retriever a good family dog?"

262 }

263 ]

264 }'

265```

266 

267

268 

269

270 

271

272Audio input to model

273 

274 Use audio inputs for prompting a model

275 

276```javascript

277import OpenAI from "openai";

278const openai = new OpenAI();

279 

280// Fetch an audio file and convert it to a base64 string

281const url = "https://cdn.openai.com/API/docs/audio/alloy.wav";

282const audioResponse = await fetch(url);

283const buffer = await audioResponse.arrayBuffer();

284const base64str = Buffer.from(buffer).toString("base64");

285 

286const response = await openai.chat.completions.create({

287 model: "gpt-audio-1.5",

288 modalities: ["text", "audio"],

289 audio: { voice: "alloy", format: "wav" },

290 messages: [

291 {

292 role: "user",

293 content: [

294 { type: "text", text: "What is in this recording?" },

295 {

296 type: "input_audio",

297 input_audio: { data: base64str, format: "wav" },

298 },

299 ],

300 },

301 ],

302 store: true,

303});

304 

305console.log(response.choices[0]);

306```

307 

308```python

309import base64

310import requests

311from openai import OpenAI

312 

313client = OpenAI()

314 

315# Fetch the audio file and convert it to a base64 encoded string

316url = "https://cdn.openai.com/API/docs/audio/alloy.wav"

317response = requests.get(url)

318response.raise_for_status()

319wav_data = response.content

320encoded_string = base64.b64encode(wav_data).decode("utf-8")

321 

322completion = client.chat.completions.create(

323 model="gpt-audio-1.5",

324 modalities=["text", "audio"],

325 audio={"voice": "alloy", "format": "wav"},

326 messages=[

327 {

328 "role": "user",

329 "content": [

330 {"type": "text", "text": "What is in this recording?"},

331 {

332 "type": "input_audio",

333 "input_audio": {"data": encoded_string, "format": "wav"},

334 },

335 ],

336 },

337 ],

338)

339 

340print(completion.choices[0].message)

341```

342 

343```go

344package main

345 

346import (

347 "context"

348 "encoding/base64"

349 "fmt"

350 "os"

351 

352 "github.com/openai/openai-go/v3"

353)

354 

355func main() {

356 audio, err := os.ReadFile("fixtures/audio.wav")

357 if err != nil {

358 panic(err)

359 }

360 client := openai.NewClient()

361 response, err := client.Chat.Completions.New(context.Background(), openai.ChatCompletionNewParams{

362 Model: "gpt-audio-1.5",

363 Modalities: []string{"text", "audio"},

364 Audio: openai.ChatCompletionAudioParam{

365 Voice: openai.ChatCompletionAudioParamVoiceUnion{OfString: openai.String("alloy")},

366 Format: openai.ChatCompletionAudioParamFormatWAV,

367 },

368 Messages: []openai.ChatCompletionMessageParamUnion{openai.UserMessage([]openai.ChatCompletionContentPartUnionParam{

369 openai.TextContentPart("What is in this recording?"),

370 openai.InputAudioContentPart(openai.ChatCompletionContentPartInputAudioInputAudioParam{

371 Data: base64.StdEncoding.EncodeToString(audio),

372 Format: "wav",

373 }),

374 })},

375 })

376 if err != nil {

377 panic(err)

378 }

379 fmt.Println(response.Choices[0])

380}

381```

382 

383```java

384import com.openai.client.OpenAIClient;

385import com.openai.client.okhttp.OpenAIOkHttpClient;

386import com.openai.models.chat.completions.ChatCompletionAudioParam;

387import com.openai.models.chat.completions.ChatCompletionContentPart;

388import com.openai.models.chat.completions.ChatCompletionContentPartInputAudio;

389import com.openai.models.chat.completions.ChatCompletionContentPartText;

390import com.openai.models.chat.completions.ChatCompletionCreateParams;

391import com.openai.models.chat.completions.ChatCompletionUserMessageParam;

392import java.io.IOException;

393import java.nio.file.Files;

394import java.nio.file.Path;

395import java.util.Base64;

396import java.util.List;

397 

398String encodedAudio =

399 Base64.getEncoder()

400 .encodeToString(

401 Files.readAllBytes(Path.of(System.getenv("OPENAI_EXAMPLE_AUDIO_PATH"))));

402 

403ChatCompletionCreateParams params =

404 ChatCompletionCreateParams.builder()

405 .model("gpt-audio-1.5")

406 .addMessage(

407 ChatCompletionUserMessageParam.builder()

408 .contentOfArrayOfContentParts(

409 List.of(

410 ChatCompletionContentPart.ofText(

411 ChatCompletionContentPartText.builder()

412 .text("What is in this recording?")

413 .build()),

414 ChatCompletionContentPart.ofInputAudio(

415 ChatCompletionContentPartInputAudio.builder()

416 .inputAudio(

417 ChatCompletionContentPartInputAudio.InputAudio.builder()

418 .data(encodedAudio)

419 .format(

420 ChatCompletionContentPartInputAudio.InputAudio

421 .Format.WAV)

422 .build())

423 .build())))

424 .build())

425 .addModality(ChatCompletionCreateParams.Modality.TEXT)

426 .addModality(ChatCompletionCreateParams.Modality.AUDIO)

427 .audio(

428 ChatCompletionAudioParam.builder()

429 .voice("alloy")

430 .format(ChatCompletionAudioParam.Format.WAV)

431 .build())

432 .store(true)

433 .build();

434 

435client.chat().completions().create(params).choices().stream()

436 .flatMap(choice -> choice.message().content().stream())

437 .forEach(System.out::println);

438```

439 

440```csharp

441using OpenAI.Chat;

442#pragma warning disable OPENAI001

443 

444string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;

445ChatClient client = new("gpt-audio-1.5", key);

446 

447BinaryData audio = BinaryData.FromBytes(

448 await File.ReadAllBytesAsync("audio.wav")

449);

450UserChatMessage message = new(

451 [

452 ChatMessageContentPart.CreateTextPart("What is in this recording?"),

453 ChatMessageContentPart.CreateInputAudioPart(

454 audio,

455 ChatInputAudioFormat.Wav

456 ),

457 ]

458);

459ChatCompletionOptions options = new()

460{

461 ResponseModalities = ChatResponseModalities.Text | ChatResponseModalities.Audio,

462 AudioOptions = new(ChatOutputAudioVoice.Alloy, ChatOutputAudioFormat.Wav),

463 StoredOutputEnabled = true,

464};

465 

466ChatCompletion completion = await client.CompleteChatAsync([message], options);

467 

468if (completion.OutputAudio is not ChatOutputAudio audioOutput)

469{

470 throw new InvalidOperationException("No audio output was returned.");

471}

472 

473Console.WriteLine(audioOutput.Transcript);

474```

475 

476```ruby

477require "base64"

478require "openai"

479 

480client = OpenAI::Client.new

481audio = Base64.strict_encode64(File.binread("audio.wav"))

482completion = client.chat.completions.create(

483 model: "gpt-audio-1.5",

484 messages: [{

485 role: :user,

486 content: [

487 {type: :text, text: "What is in this recording?"},

488 {type: :input_audio, input_audio: {data: audio, format: :wav}}

489 ]

490 }],

491 modalities: [:text, :audio],

492 audio: {voice: :alloy, format: :wav},

493 store: true

494)

495 

496puts(completion.choices.fetch(0).message.content)

497```

498 

499```bash

500curl "https://api.openai.com/v1/chat/completions" \

501 -H "Content-Type: application/json" \

502 -H "Authorization: Bearer $OPENAI_API_KEY" \

503 -d '{

504 "model": "gpt-audio-1.5",

505 "modalities": ["text", "audio"],

506 "audio": { "voice": "alloy", "format": "wav" },

507 "messages": [

508 {

509 "role": "user",

510 "content": [

511 { "type": "text", "text": "What is in this recording?" },

512 {

513 "type": "input_audio",

514 "input_audio": {

515 "data": "<base64 bytes here>",

516 "format": "wav"

517 }

518 }

519 ]

520 }

521 ]

522 }'

523```

Details

1# Audio in Chat Completions

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5If you already have a text-based LLM application with the [Chat Completions endpoint](https://developers.openai.com/api/reference/resources/chat), you may want to add audio capabilities. For example, if your chat application supports text input, you can add audio input and output: include `audio` in the `modalities` array and use an audio model, like [`gpt-audio-1.5`](https://developers.openai.com/api/docs/models/gpt-audio-1.5).

6 

7The [Responses API](https://developers.openai.com/api/reference/resources/responses) docs currently describe

8 text and image inputs with text outputs. For this audio-chat pattern, use Chat

9 Completions with an audio-capable model.

10 

11 

12 

13Audio output from model

14 

15 Create a human-like audio response to a prompt

16 

17```javascript

18import { writeFileSync } from "node:fs";

19import OpenAI from "openai";

20 

21const openai = new OpenAI();

22 

23// Generate an audio response to the given prompt

24const response = await openai.chat.completions.create({

25 model: "gpt-audio-1.5",

26 modalities: ["text", "audio"],

27 audio: { voice: "alloy", format: "wav" },

28 messages: [

29 {

30 role: "user",

31 content: "Is a golden retriever a good family dog?",

32 },

33 ],

34 store: true,

35});

36 

37// Inspect returned data

38console.log(response.choices[0]);

39 

40// Write audio data to a file

41writeFileSync(

42 "dog.wav",

43 Buffer.from(response.choices[0].message.audio.data, "base64"),

44 { encoding: "utf-8" }

45);

46```

47 

48```python

49import base64

50from openai import OpenAI

51 

52client = OpenAI()

53 

54completion = client.chat.completions.create(

55 model="gpt-audio-1.5",

56 modalities=["text", "audio"],

57 audio={"voice": "alloy", "format": "wav"},

58 messages=[{"role": "user", "content": "Is a golden retriever a good family dog?"}],

59)

60 

61print(completion.choices[0])

62 

63wav_bytes = base64.b64decode(completion.choices[0].message.audio.data)

64with open("dog.wav", "wb") as f:

65 f.write(wav_bytes)

66```

67 

68```go

69package main

70 

71import (

72 "context"

73 "encoding/base64"

74 "fmt"

75 "os"

76 

77 "github.com/openai/openai-go/v3"

78)

79 

80func main() {

81 client := openai.NewClient()

82 response, err := client.Chat.Completions.New(context.Background(), openai.ChatCompletionNewParams{

83 Model: "gpt-audio-1.5",

84 Modalities: []string{"text", "audio"},

85 Audio: openai.ChatCompletionAudioParam{

86 Voice: openai.ChatCompletionAudioParamVoiceUnion{OfString: openai.String("alloy")},

87 Format: openai.ChatCompletionAudioParamFormatWAV,

88 },

89 Messages: []openai.ChatCompletionMessageParamUnion{openai.UserMessage("Is a golden retriever a good family dog?")},

90 })

91 if err != nil {

92 panic(err)

93 }

94 fmt.Println(response.Choices[0])

95 audio, err := base64.StdEncoding.DecodeString(response.Choices[0].Message.Audio.Data)

96 if err != nil {

97 panic(err)

98 }

99 if err := os.WriteFile("dog.wav", audio, 0o600); err != nil {

100 panic(err)

101 }

102}

103```

104 

105```java

106import com.openai.client.OpenAIClient;

107import com.openai.client.okhttp.OpenAIOkHttpClient;

108import com.openai.models.chat.completions.ChatCompletionAudioParam;

109import com.openai.models.chat.completions.ChatCompletionCreateParams;

110import java.io.IOException;

111import java.nio.file.Files;

112import java.nio.file.Path;

113import java.util.Base64;

114 

115ChatCompletionCreateParams params =

116 ChatCompletionCreateParams.builder()

117 .model("gpt-audio-1.5")

118 .addUserMessage("Is a golden retriever a good family dog?")

119 .addModality(ChatCompletionCreateParams.Modality.TEXT)

120 .addModality(ChatCompletionCreateParams.Modality.AUDIO)

121 .audio(

122 ChatCompletionAudioParam.builder()

123 .voice("alloy")

124 .format(ChatCompletionAudioParam.Format.WAV)

125 .build())

126 .store(true)

127 .build();

128 

129var message = client.chat().completions().create(params).choices().get(0).message();

130var audio =

131 message.audio().orElseThrow(() -> new IllegalStateException("No audio output returned"));

132Files.write(Path.of("dog.wav"), Base64.getDecoder().decode(audio.data()));

133message.content().ifPresent(System.out::println);

134```

135 

136```csharp

137using OpenAI.Chat;

138#pragma warning disable OPENAI001

139 

140string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;

141ChatClient client = new("gpt-audio-1.5", key);

142 

143ChatCompletionOptions options = new()

144{

145 ResponseModalities = ChatResponseModalities.Text | ChatResponseModalities.Audio,

146 AudioOptions = new(ChatOutputAudioVoice.Alloy, ChatOutputAudioFormat.Wav),

147 StoredOutputEnabled = true,

148};

149 

150ChatCompletion completion = await client.CompleteChatAsync(

151 [new UserChatMessage("Is a golden retriever a good family dog?")],

152 options

153);

154 

155if (completion.OutputAudio is not ChatOutputAudio audio)

156{

157 throw new InvalidOperationException("No audio output was returned.");

158}

159 

160Console.WriteLine(audio.Transcript);

161await File.WriteAllBytesAsync("dog.wav", audio.AudioBytes.ToArray());

162```

163 

164```ruby

165require "base64"

166require "openai"

167 

168client = OpenAI::Client.new

169completion = client.chat.completions.create(

170 model: "gpt-audio-1.5",

171 messages: [{role: :user, content: "Is a golden retriever a good family dog?"}],

172 modalities: [:text, :audio],

173 audio: {voice: :alloy, format: :wav},

174 store: true

175)

176 

177audio = completion.choices.fetch(0).message.audio or raise "No audio returned"

178File.binwrite("dog.wav", Base64.strict_decode64(audio.data))

179```

180 

181```bash

182curl "https://api.openai.com/v1/chat/completions" \

183 -H "Content-Type: application/json" \

184 -H "Authorization: Bearer $OPENAI_API_KEY" \

185 -d '{

186 "model": "gpt-audio-1.5",

187 "modalities": ["text", "audio"],

188 "audio": { "voice": "alloy", "format": "wav" },

189 "messages": [

190 {

191 "role": "user",

192 "content": "Is a golden retriever a good family dog?"

193 }

194 ]

195 }'

196```

197 

198

199 

200

201 

202

203Audio input to model

204 

205 Use audio inputs for prompting a model

206 

207```javascript

208import OpenAI from "openai";

209const openai = new OpenAI();

210 

211// Fetch an audio file and convert it to a base64 string

212const url = "https://cdn.openai.com/API/docs/audio/alloy.wav";

213const audioResponse = await fetch(url);

214const buffer = await audioResponse.arrayBuffer();

215const base64str = Buffer.from(buffer).toString("base64");

216 

217const response = await openai.chat.completions.create({

218 model: "gpt-audio-1.5",

219 modalities: ["text", "audio"],

220 audio: { voice: "alloy", format: "wav" },

221 messages: [

222 {

223 role: "user",

224 content: [

225 { type: "text", text: "What is in this recording?" },

226 {

227 type: "input_audio",

228 input_audio: { data: base64str, format: "wav" },

229 },

230 ],

231 },

232 ],

233 store: true,

234});

235 

236console.log(response.choices[0]);

237```

238 

239```python

240import base64

241import requests

242from openai import OpenAI

243 

244client = OpenAI()

245 

246# Fetch the audio file and convert it to a base64 encoded string

247url = "https://cdn.openai.com/API/docs/audio/alloy.wav"

248response = requests.get(url)

249response.raise_for_status()

250wav_data = response.content

251encoded_string = base64.b64encode(wav_data).decode("utf-8")

252 

253completion = client.chat.completions.create(

254 model="gpt-audio-1.5",

255 modalities=["text", "audio"],

256 audio={"voice": "alloy", "format": "wav"},

257 messages=[

258 {

259 "role": "user",

260 "content": [

261 {"type": "text", "text": "What is in this recording?"},

262 {

263 "type": "input_audio",

264 "input_audio": {"data": encoded_string, "format": "wav"},

265 },

266 ],

267 },

268 ],

269)

270 

271print(completion.choices[0].message)

272```

273 

274```go

275package main

276 

277import (

278 "context"

279 "encoding/base64"

280 "fmt"

281 "os"

282 

283 "github.com/openai/openai-go/v3"

284)

285 

286func main() {

287 audio, err := os.ReadFile("fixtures/audio.wav")

288 if err != nil {

289 panic(err)

290 }

291 client := openai.NewClient()

292 response, err := client.Chat.Completions.New(context.Background(), openai.ChatCompletionNewParams{

293 Model: "gpt-audio-1.5",

294 Modalities: []string{"text", "audio"},

295 Audio: openai.ChatCompletionAudioParam{

296 Voice: openai.ChatCompletionAudioParamVoiceUnion{OfString: openai.String("alloy")},

297 Format: openai.ChatCompletionAudioParamFormatWAV,

298 },

299 Messages: []openai.ChatCompletionMessageParamUnion{openai.UserMessage([]openai.ChatCompletionContentPartUnionParam{

300 openai.TextContentPart("What is in this recording?"),

301 openai.InputAudioContentPart(openai.ChatCompletionContentPartInputAudioInputAudioParam{

302 Data: base64.StdEncoding.EncodeToString(audio),

303 Format: "wav",

304 }),

305 })},

306 })

307 if err != nil {

308 panic(err)

309 }

310 fmt.Println(response.Choices[0])

311}

312```

313 

314```java

315import com.openai.client.OpenAIClient;

316import com.openai.client.okhttp.OpenAIOkHttpClient;

317import com.openai.models.chat.completions.ChatCompletionAudioParam;

318import com.openai.models.chat.completions.ChatCompletionContentPart;

319import com.openai.models.chat.completions.ChatCompletionContentPartInputAudio;

320import com.openai.models.chat.completions.ChatCompletionContentPartText;

321import com.openai.models.chat.completions.ChatCompletionCreateParams;

322import com.openai.models.chat.completions.ChatCompletionUserMessageParam;

323import java.io.IOException;

324import java.nio.file.Files;

325import java.nio.file.Path;

326import java.util.Base64;

327import java.util.List;

328 

329String encodedAudio =

330 Base64.getEncoder()

331 .encodeToString(

332 Files.readAllBytes(Path.of(System.getenv("OPENAI_EXAMPLE_AUDIO_PATH"))));

333 

334ChatCompletionCreateParams params =

335 ChatCompletionCreateParams.builder()

336 .model("gpt-audio-1.5")

337 .addMessage(

338 ChatCompletionUserMessageParam.builder()

339 .contentOfArrayOfContentParts(

340 List.of(

341 ChatCompletionContentPart.ofText(

342 ChatCompletionContentPartText.builder()

343 .text("What is in this recording?")

344 .build()),

345 ChatCompletionContentPart.ofInputAudio(

346 ChatCompletionContentPartInputAudio.builder()

347 .inputAudio(

348 ChatCompletionContentPartInputAudio.InputAudio.builder()

349 .data(encodedAudio)

350 .format(

351 ChatCompletionContentPartInputAudio.InputAudio

352 .Format.WAV)

353 .build())

354 .build())))

355 .build())

356 .addModality(ChatCompletionCreateParams.Modality.TEXT)

357 .addModality(ChatCompletionCreateParams.Modality.AUDIO)

358 .audio(

359 ChatCompletionAudioParam.builder()

360 .voice("alloy")

361 .format(ChatCompletionAudioParam.Format.WAV)

362 .build())

363 .store(true)

364 .build();

365 

366client.chat().completions().create(params).choices().stream()

367 .flatMap(choice -> choice.message().content().stream())

368 .forEach(System.out::println);

369```

370 

371```csharp

372using OpenAI.Chat;

373#pragma warning disable OPENAI001

374 

375string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;

376ChatClient client = new("gpt-audio-1.5", key);

377 

378BinaryData audio = BinaryData.FromBytes(

379 await File.ReadAllBytesAsync("audio.wav")

380);

381UserChatMessage message = new(

382 [

383 ChatMessageContentPart.CreateTextPart("What is in this recording?"),

384 ChatMessageContentPart.CreateInputAudioPart(

385 audio,

386 ChatInputAudioFormat.Wav

387 ),

388 ]

389);

390ChatCompletionOptions options = new()

391{

392 ResponseModalities = ChatResponseModalities.Text | ChatResponseModalities.Audio,

393 AudioOptions = new(ChatOutputAudioVoice.Alloy, ChatOutputAudioFormat.Wav),

394 StoredOutputEnabled = true,

395};

396 

397ChatCompletion completion = await client.CompleteChatAsync([message], options);

398 

399if (completion.OutputAudio is not ChatOutputAudio audioOutput)

400{

401 throw new InvalidOperationException("No audio output was returned.");

402}

403 

404Console.WriteLine(audioOutput.Transcript);

405```

406 

407```ruby

408require "base64"

409require "openai"

410 

411client = OpenAI::Client.new

412audio = Base64.strict_encode64(File.binread("audio.wav"))

413completion = client.chat.completions.create(

414 model: "gpt-audio-1.5",

415 messages: [{

416 role: :user,

417 content: [

418 {type: :text, text: "What is in this recording?"},

419 {type: :input_audio, input_audio: {data: audio, format: :wav}}

420 ]

421 }],

422 modalities: [:text, :audio],

423 audio: {voice: :alloy, format: :wav},

424 store: true

425)

426 

427puts(completion.choices.fetch(0).message.content)

428```

429 

430```bash

431curl "https://api.openai.com/v1/chat/completions" \

432 -H "Content-Type: application/json" \

433 -H "Authorization: Bearer $OPENAI_API_KEY" \

434 -d '{

435 "model": "gpt-audio-1.5",

436 "modalities": ["text", "audio"],

437 "audio": { "voice": "alloy", "format": "wav" },

438 "messages": [

439 {

440 "role": "user",

441 "content": [

442 { "type": "text", "text": "What is in this recording?" },

443 {

444 "type": "input_audio",

445 "input_audio": {

446 "data": "<base64 bytes here>",

447 "format": "wav"

448 }

449 }

450 ]

451 }

452 ]

453 }'

454```

guides/custom-voices.md +225 −0 created

Details

1# Custom voices

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5Custom voices enable you to create a unique voice for your agent or application. These voices can be used for audio output with the [Text to Speech API](https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create), the [Realtime API](https://developers.openai.com/api/reference/resources/realtime), or the [Chat Completions API with audio output](https://developers.openai.com/api/docs/guides/audio-chat-completions).

6 

7To create a custom voice, you’ll provide a short sample audio reference that the model will seek to replicate.

8 

9 

10 

11 {"Custom voices are limited to eligible customers. Contact our "}

12 [{"sales team"}](https://openai.com/contact-sales/)

13 {

14 " to learn more. Once enabled for your organization, you’ll have access to the "

15 }

16 [{"Voices"}](https://platform.openai.com/audio/voices)

17 {" tab under Audio."}

18

19 

20 

21## Creating a voice

22 

23Currently, voices must be created through an API request. See the API reference for the full set of API operations.

24 

25Creating a voice requires two separate audio recordings:

26 

271. **Consent recording:** This recording captures the voice actor providing consent to create a likeness of their voice. The actor must read one of the consent phrases provided below.

282. **Sample recording:** The actual audio sample that the model will try to adhere to. The voice must match the consent recording.

29 

30**Tips for creating a high-quality voice**

31 

32The quality of your custom voice is highly dependent on the quality of the sample you provide. Optimizing the recording quality can make a big difference.

33 

34- Record in a quiet space with minimal echo.

35- Use a professional XLR microphone.

36- Stay about 7–8 inches from the mic with a pop filter in between, and keep that distance consistent.

37- The model copies exactly what you give it—tone, cadence, energy, pauses, habits—so record the exact voice you want. Be consistent in energy, style, and accent throughout.

38- Small variations in the audio sample can result in quality differences with the generated voice. Try multiple examples to find the best fit.

39 

40**Requirements and limitations**

41 

42- At most 20 voices can be created per organization.

43- The audio samples must be 30 seconds or less.

44- The audio samples must be one of the following types: `mpeg`, `wav`, `ogg`, `aac`, `flac`, `webm`, or `mp4`.

45 

46Refer to the Text-to-Speech Supplemental Agreement for additional terms of use.

47 

48**Creating a voice consent**

49 

50The consent audio recording must only include one of the following phrases. Any divergence from the script will lead to a failure.

51 

52| Language | Phrase |

53| -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |

54| `de` | Ich bin der Eigentümer dieser Stimme und bin damit einverstanden, dass OpenAI diese Stimme zur Erstellung eines synthetischen Stimmmodells verwendet. |

55| `en` | I am the owner of this voice and I consent to OpenAI using this voice to create a synthetic voice model. |

56| `es` | Soy el propietario de esta voz y doy mi consentimiento para que OpenAI la utilice para crear un modelo de voz sintética. |

57| `fr` | Je suis le propriétaire de cette voix et j'autorise OpenAI à utiliser cette voix pour créer un modèle de voix synthétique. |

58| `hi` | मैं इस आवाज का मालिक हूं और मैं सिंथेटिक आवाज मॉडल बनाने के लिए OpenAI को इस आवाज का उपयोग करने की सहमति देता हूं |

59| `id` | Saya adalah pemilik suara ini dan saya memberikan persetujuan kepada OpenAI untuk menggunakan suara ini guna membuat model suara sintetis. |

60| `it` | Sono il proprietario di questa voce e acconsento che OpenAI la utilizzi per creare un modello di voce sintetica. |

61| `ja` | 私はこの音声の所有者であり、OpenAIがこの音声を使用して音声合成 モデルを作成することを承認します。 |

62| `ko` | 나는 이 음성의 소유자이며 OpenAI가 이 음성을 사용하여 음성 합성 모델을 생성할 것을 허용합니다. |

63| `nl` | Ik ben de eigenaar van deze stem en ik geef OpenAI toestemming om deze stem te gebruiken om een synthetisch stemmodel te maken. |

64| `pl` | Jestem właścicielem tego głosu i wyrażam zgodę na wykorzystanie go przez OpenAI w celu utworzenia syntetycznego modelu głosu. |

65| `pt` | Eu sou o proprietário desta voz e autorizo o OpenAI a usá-la para criar um modelo de voz sintética. |

66| `ru` | Я являюсь владельцем этого голоса и даю согласие OpenAI на использование этого голоса для создания модели синтетического голоса. |

67| `uk` | Я є власником цього голосу і даю згоду OpenAI використовувати цей голос для створення синтетичної голосової моделі. |

68| `vi` | Tôi là chủ sở hữu giọng nói này và tôi đồng ý cho OpenAI sử dụng giọng nói này để tạo mô hình giọng nói tổng hợp. |

69| `zh` | 我是此声音的拥有者并授权OpenAI使用此声音创建语音合成模型 |

70 

71Then upload the recording via the API. A successful upload will return the consent recording ID that you’ll reference later. Note the consent can be used for multiple different voice creations if the same voice actor is making multiple attempts.

72 

73```bash

74curl https://api.openai.com/v1/audio/voice_consents \

75 -X POST \

76 -H "Authorization: Bearer $OPENAI_API_KEY" \

77 -F "name=test_consent" \

78 -F "language=en" \

79 -F "recording=@$HOME/tmp/voice_consent/consent_recording.wav;type=audio/x-wav"

80```

81 

82 

83**Creating a voice**

84 

85Next, you’ll create the actual voice by referencing the consent recording ID, and providing the voice sample.

86 

87```bash

88curl https://api.openai.com/v1/audio/voices \

89 -X POST \

90 -H "Authorization: Bearer $OPENAI_API_KEY" \

91 -F "name=test_voice" \

92 -F "audio_sample=@$HOME/tmp/voice_consent/audio_sample_recording.wav;type=audio/x-wav" \

93 -F "consent=cons_123abc"

94```

95 

96 

97If successful, the created voice will be listed under the [Audio tab](https://platform.openai.com/audio/voices).

98 

99## Using a voice during speech generation

100 

101Speech generation will work as usual. Specify the ID of the voice in the `voice` parameter when [creating speech](https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create), or when initiating a [realtime session](https://developers.openai.com/api/reference/resources/realtime/subresources/calls/methods/create#realtime_create_call-session-audio-output-voice).

102 

103**Text to speech example**

104 

105```bash

106curl https://api.openai.com/v1/audio/speech \

107 -X POST \

108 -H "Authorization: Bearer $OPENAI_API_KEY" \

109 -H "Content-Type: application/json" \

110 -d '{

111 "model": "gpt-4o-mini-tts",

112 "voice": {

113 "id": "voice_123abc"

114 },

115 "input": "Maple est le meilleur golden retriever du monde entier.",

116 "language": "fr",

117 "format": "wav"

118 }' \

119 --output sample.wav

120```

121 

122 

123**Realtime API example**

124 

125For Ruby, set `OPENAI_VOICE_ID` to your custom voice ID before running the example.

126 

127```javascript

128const sessionConfig = JSON.stringify({

129 session: {

130 type: "realtime",

131 model: "gpt-realtime-2",

132 audio: {

133 output: {

134 voice: { id: "voice_123abc" },

135 },

136 },

137 },

138});

139```

140 

141```ruby

142require "json"

143 

144session_config = JSON.generate(

145 session: {

146 type: "realtime",

147 model: "gpt-realtime-2",

148 audio: {output: {voice: {id: ENV.fetch("OPENAI_VOICE_ID")}}}

149 }

150)

151puts(session_config)

152```

153 

154 

155## Use a custom voice with GPT-Live

156 

157Use a project-scoped API key approved for both GPT-Live and custom voice

158creation. Reading consent phrases and using a custom voice require

159`api.voices.read`; creating consents and voices requires `api.voices.write` and

160custom-voice API access. Use the same project for every request and keep the API

161key on a trusted server.

162 

163### Prepare the recordings

164 

165List the current supported consent phrases before recording:

166 

167```bash

168curl https://api.openai.com/v1/audio/consent_phrases \

169 -H "Authorization: Bearer $OPENAI_API_KEY"

170```

171 

172The consent recording and reference sample must come from the same person. The

173sample needs at least five seconds of actual speech and at least 15 transcribed

174text tokens; silence does not count. Use a 10–30-second

175recording with several complete sentences. Each upload is limited to 10 MiB.

176The service extracts the reference transcript; do not upload transcript tokens,

177configure a decoder, or add custom request headers.

178 

179Browser recorders may label audio `audio/webm;codecs=opus`, which the upload

180endpoint rejects. When constructing an upload, use the supported base MIME type

181`audio/webm` while preserving the original audio bytes. Use the consent and voice

182creation requests above, then save the returned voice ID.

183 

184### Select the voice at session creation

185 

186Pass a custom voice as the object `{ "id": "voice_123" }`, not the string

187`"voice_123"`. Named voices such as `"marin"` use strings.

188 

189`gpt-live-1` supports custom voices with English accents. To use an accent, also

190specify it in `session.instructions`, such as "Speak British English" or "Speak

191Irish English." The example below uses British English; change the instruction

192to match the accent you want for your custom voice.

193 

194Include the following configuration in the initial session:

195 

196```json

197{

198 "model": "gpt-live-1",

199 "instructions": "You are a helpful voice assistant. Speak British English.",

200 "audio": { "output": { "voice": { "id": "voice_123" } } }

201}

202```

203 

204For [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live), the trusted session broker

205places this configuration in the JSON `session` field alongside `transport`.

206The Live endpoint requires JSON, not multipart or raw SDP. Read the created

207session ID from `session.id` and the SDP answer from `transport.sdp`. Authenticate

208hosted broker requests with application credentials; never expose the OpenAI

209API key to the browser.

210 

211For [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live), put the configuration

212in the first `session.start` event. Connect without query parameters and wait

213for `session.started` before streaming audio. Send audio with

214`session.input_audio.append`. After sending `session.close`, keep receiving until

215`session.closed` supplies final usage.

216 

217### Handle access and lifecycle failures

218 

219- The output voice cannot be changed after the Live session starts. Start a new session to use a different voice.

220- A deleted or revoked voice, a consent from another project, or missing custom-voice access can appear as a `404`.

221- Malformed audio, a mismatched speaker, or a non-project-scoped key is rejected.

222 

223Confirm your project's permissions, recording minimums, and upload limits

224before creating a voice. See [GPT-Live getting started](https://developers.openai.com/api/docs/guides/live)

225for session setup requirements.

guides/live.md +58 −0 created

Details

1# Getting started with GPT-Live

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5GPT-Live handles a spoken conversation while a backend agent looks up information, uses tools, and completes tasks. It can listen while speaking, a capability called **full duplex**. Sending work to the backend is called **delegation**: the conversation can continue while that work runs.

6 

7For example, a user can ask about an order, add a detail while the backend checks its status, and hear the result when it is ready. You choose the backend model or agent independently of the voice model; Realtime uses one model for speech, reasoning, and tool selection.

8 

9## Understand the two parts

10 

11- **GPT-Live handles conversation.** It listens, speaks, and decides when to ask the backend for help. Give it a short prompt for conversation style and when to delegate.

12- **The backend handles delegated tasks.** With Responses delegation, use a supported Responses model. With client delegation, connect any model, agent harness, or service your application runs. The backend reasons, uses tools, and returns results for GPT-Live to communicate. Keep detailed instructions, business rules, and tool workflows here.

13 

14Your application owns permissions, confirmations, private function execution, and durable task state. Interrupting speech does not automatically cancel backend work. See [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) to compare GPT-Live with Realtime and chained voice applications.

15 

16 

17 

18 

19 

20## Choose how to run the backend

21 

22Start with **[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#configure-responses-delegation)** when a managed backend fits: GPT-Live calls your configured Responses model, supplies conversation context, and returns results. Choose **[client delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#configure-client-delegation)** when your application needs to control backend execution, context, or which results reach GPT-Live.

23 

24See [Choose a delegation mode](https://developers.openai.com/api/docs/guides/live-delegation#choose-a-delegation-mode) for the comparison and configuration details. Choose the mode when you create the session; to change modes, start a new session.

25 

26## Connect your first session

27 

28Start with the [GPT-Live WebRTC quickstart](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live). Its browser and server example connects microphone input and speaker output, with a Responses backend that can search the web.

29 

30You need a microphone, a browser page served over HTTPS or localhost, and a trusted server with an OpenAI project API key. Keep the key on the server.

31 

321. Write a short [conversation prompt](https://developers.openai.com/api/docs/guides/live-prompting) that tells GPT-Live when to ask the backend for help.

332. Follow the quickstart to connect the browser's microphone, audio playback, and event channel. Your server creates the session and exchanges the browser's connection offer for an answer.

343. Wait for `session.started`, then speak and listen to a reply. Ask a question that needs current information to try the web search backend.

354. End the conversation and [close the session](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) to collect final usage and release the connection.

36 

37Check both the spoken conversation and the backend result. A session-start event confirms startup; listening to a reply and checking the search result verify separate parts of the application.

38 

39GPT-Live voice sessions are billed by duration, per second. See the [model pricing](https://developers.openai.com/api/docs/models/gpt-live-1) for the current rate. Backend model and tool usage is billed separately. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting and ways to reduce costs.

40 

41## Choose a connection

42 

43- **[WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live)** for browser voice applications. Media tracks carry audio; a data channel carries JSON events.

44- **[WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live)** for server-side audio integrations. The primary socket carries audio and control events.

45- **[Server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live)** for backend access to an existing session. A sideband connection carries events while audio stays on the primary connection.

46- **[Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip?api=live)** for phone integration paths and provider guidance.

47 

48## Partner integrations

49 

50For applications built with **LiveKit**, **Twilio**, **Telnyx**, or **Daily/Pipecat**, start with the [partner integration overview](https://developers.openai.com/api/docs/guides/live-partner-integrations) to choose a connection for your existing media path.

51 

52## Continue building

53 

54- Shape conversation style and delegation behavior in [Prompting GPT-Live](https://developers.openai.com/api/docs/guides/live-prompting).

55- Connect tools and reduce backend latency in [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation).

56- Manage context, transcripts, and session lifecycle in [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).

57- Choose a migration path for your Realtime or text-based agent in [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration).

58- Test conversation and task outcomes in [Evaluating voice agents](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation).

guides/live-conversations.md +575 −0 created

Details

1# Managing GPT-Live sessions

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5After [connecting to GPT-Live](https://developers.openai.com/api/docs/guides/live), use session events to update context, display transcripts, and manage the connection's lifecycle. The model can listen and speak at the same time, so keep received events, audio playback, and backend task state separate in your application.

6 

7This guide assumes your connection has emitted `session.started`. See [Connections](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for connection setup and audio streaming, and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for backend work.

8 

9 

10 

11 

12 

13## Configure a session

14 

15Choose the model, voice, and delegation mode when you create the session. Give the model instructions for the conversation and include relevant history. GPT-Live manages context automatically as the conversation grows.

16 

17### Configuration fields

18 

19| Setting | Configure at startup | Change during the session |

20| ------------ | --------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |

21| Model | Set the required `model`. | Start a new session to change it. |

22| Instructions | Set `instructions` for conversation behavior, up to 16,384 tokens. | Add instructions with `session.instructions.append`. |

23| History | Set `input` to relevant prior text messages. It defaults to `[]`. | Append context; don't replace the startup history. |

24| Voice | Set `audio.output.voice` to a supported voice or authorized custom voice. The default is `marin`. | Start a new session to change it. |

25| Delegation | Set `delegation.type` to `client` or `responses`. Omitted or `null` delegation selects client mode. | Update Responses settings within the existing mode. |

26| Storage | Set `store` to `true` to make the session available for forking. It defaults to `false`. | Choose at startup. |

27 

28### Voice options

29 

30Choose a voice when you create the session. Set `audio.output.voice` to the API name, such as `"quartz"`. GPT-Live includes these additional voice options:

31 

32| Voice | API name | Language | Regional influence | Presentation | Source |

33| -------- | ---------- | ---------- | ------------------ | ------------ | --------- |

34| Quartz | `quartz` | English | Australian | Feminine | Generated |

35| Ripple | `ripple` | English | Australian | Masculine | Natural |

36| Vesper | `vesper` | English | British | Masculine | Natural |

37| Willow | `willow` | English | Irish | Feminine | Natural |

38| Stone | `stone` | English | Irish | Masculine | Natural |

39| Gleam | `gleam` | English | North American | Feminine | Natural |

40| Meridian | `meridian` | English | North American | Masculine | Natural |

41| Bossa | `bossa` | Portuguese | Brazilian | Feminine | Natural |

42| Tempo | `tempo` | Portuguese | Brazilian | Masculine | Natural |

43| Beacon | `beacon` | English | Filipino | Masculine | Generated |

44| Delta | `delta` | English | Southern U.S. | Feminine | Generated |

45| Cinder | `cinder` | English | Southern U.S. | Masculine | Generated |

46 

47Regional influence describes a voice's speaking style, not a guarantee of accent fidelity. For an approved voice created from your own recording, see [Custom voices](https://developers.openai.com/api/docs/guides/custom-voices).

48 

49 

50 

51 

52 

53For WebSocket, choose the shared `audio.format` at startup; it cannot change during the session. For WebRTC, omit this field because the connection negotiates its audio format. See [WebSocket audio formats](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for format and streaming details.

54 

55### Update a live session

56 

57Use `session.update` for changes to `session.delegation.responses` in a session already using Responses delegation. Send only the settings you want to change; omitted settings retain their values. See [Configure Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for the settings and update workflow.

58 

59You cannot change the delegation mode after startup. In particular, setting `delegation` to `null` selects client mode; it does not reset a Responses session. The startup fields `model`, `instructions`, `input`, `audio`, and `store` are not accepted update fields. Unknown configuration fields are rejected.

60 

61A successful update emits `session.updated` with the full resolved session configuration. When you supply an `event_id`, the acknowledgment returns it as `client_event_id`. Check for [rejected commands](#handle-rejected-commands) as well as acknowledgments. Acceptance confirms the configuration update; it does not establish that a backend task ran or that the model spoke.

62 

63## Provide history and context

64 

65Use startup history to resume a topic, and append relevant context as the conversation continues. Keep trusted application instructions separate from user messages and factual results.

66 

67### Seed a session with prior conversation

68 

69Include prior text messages in `session.input` when you create the session. For example, add this `input` field to your [session creation configuration](https://developers.openai.com/api/docs/guides/live#connect-your-first-session):

70 

71```javascript

72/** @type {import("openai/resources/live/live").SessionConfig} */

73```

74 

75```python

76from openai.types.live.session_config_param import SessionConfigParam

77 

78session: SessionConfigParam = {

79 "model": "gpt-live-1",

80 "input": [

81 {

82 "type": "message",

83 "role": "user",

84 "content": [

85 {"type": "input_text", "text": "I need help with my recent order."}

86 ],

87 },

88 {

89 "type": "message",

90 "role": "assistant",

91 "content": [{"type": "output_text", "text": "What is the order number?"}],

92 },

93 ],

94}

95```

96 

97 

98The list accepts up to 128 messages and 8,192 combined tokens. Supported roles are `developer`, `user`, and `assistant`, each with one text part. Developer and user messages use `input_text`; assistant messages use `text` or `output_text`. Put trusted application instructions in `instructions` or a developer message. The list does not accept the `system` role.

99 

100Select the history needed for the next interaction. `input` is a startup field, not a way to replace history during a running session. It also does not accept the full range of backend input items used in Responses delegation.

101 

102### Understand when context reaches the model

103 

104The full `input` supplied at session creation is available to the model when the session starts. Put context the model needs from the beginning in this field.

105 

106During a running session, the `session.instructions.append`, `session.thinking.append`, and `session.commentary.append` events feed content into the model over time. Their acknowledgments wait until frame progress reaches the estimated end of context injection. The returned `start_ms` and `end_ms` describe an estimated range on the session timeline, not speech or playback completion. They do not prove that the model consumed the entire update. Don't assume its next speech will reflect the whole update.

107 

108If frame progress stops, an acknowledgment can remain pending. Closing the session reports an error for pending appends. Match each acknowledgment to the outgoing `event_id` through `client_event_id`, and keep handling errors while you wait.

109 

110### Add context during the conversation

111 

112Choose an event based on how the model should use the update:

113 

114- `session.instructions.append`: add trusted application instructions that influence behavior and speech.

115- `session.thinking.append`: add factual context without asking the model to say it immediately.

116- `session.commentary.append`: provide information for the model to say aloud, which it may paraphrase.

117 

118Each event takes plain-string `content` of up to 500 tokens and a required `delegation_id`. Use `null` for session-wide context. For example, send this after your application has verified the user's acceptance and started the lookup:

119 

120```javascript

121/**

122 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

123 */

124export function sendUpdate(connection) {

125 connection.send({

126 type: "session.thinking.append",

127 event_id: "context_1",

128 delegation_id: null,

129 content:

130 "The user has already accepted the terms. The account lookup is still running.",

131 });

132}

133```

134 

135```python

136from openai.resources.live.live import AsyncLiveConnection

137from openai.resources.live.sideband import AsyncSidebandConnection

138 

139 

140async def send_update(

141 connection: AsyncLiveConnection | AsyncSidebandConnection,

142) -> None:

143 await connection.session.thinking.append(

144 event_id="context_1",

145 delegation_id=None,

146 content=(

147 "The user has already accepted the terms. The account lookup is still "

148 "running."

149 ),

150 )

151```

152 

153 

154Wait for `session.thinking.appended` with `client_event_id: "context_1"`, or handle an error. The acknowledgment confirms that context was accepted. It does not confirm speech, playback, or completion of an external action.

155 

156Quiet context can influence later speech; it is not a privacy boundary. Keep credentials, secrets, and text the model must never reveal out of all three events. Use the instructions event for application-authored behavior, not untrusted tool output. Enforce permissions and required confirmations in your application.

157 

158For page navigation, selections, and other UI changes, see [Share UI context](https://developers.openai.com/api/docs/guides/live-delegation#share-ui-context) for concise updates that help GPT-Live understand what the user is referring to.

159 

160For results tied to a backend task, use a known client delegation ID and follow [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update). That ID is not a Responses response ID or tool call ID.

161 

162Use instructions to steer the conversation after an application check triggers. Your server can monitor events and send these corrections through a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) attached to the existing session, or through its primary WebSocket. See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for concurrent checks, action blocking, and playback control.

163 

164 

165 

166 

167 

168 

169 

170 

171## Build the conversation interface

172 

173Display transcripts and microphone state independently of backend progress. Receiving assistant text does not tell you how much audio the user has heard.

174 

175### Transcript deltas

176 

177Listen for `session.input_transcript.delta` for user speech and `session.output_transcript.delta` for assistant speech. Each event contains a text fragment and its interval on the session timeline:

178 

179```json

180{

181 "type": "session.input_transcript.delta",

182 "event_id": "event_transcript_1",

183 "delta": "What is",

184 "start_ms": 1000,

185 "end_ms": 1200

186}

187```

188 

189Append fragments in order for each speaker, retaining `start_ms` and `end_ms`. These are milliseconds on the session timeline, with intervals that include the start and exclude the end. They are not wall-clock timestamps, packet arrival times, or exact word alignments.

190 

191Only intervals containing transcript text produce events, and network delivery can be uneven. Do not infer silence from a missing event or treat a fragment as a complete user turn. Transcript deltas have no item ID or authoritative turn-completed event.

192 

193Processing transcript fragments is optional. You can use them to update your UI, run checks, or start work early while the conversation continues. For lightweight checks, consider a small model such as `gpt-5.6-luna` with low reasoning effort. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples and connection guidance.

194 

195For conversation guardrails, check accumulated user and assistant text as it arrives. Transcript delivery does not provide an advance buffer for approving speech before playback. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).

196 

197 

198 

199 

200 

201If your interface groups text into turns, keep that grouping revisable. Preserve the original fragments, allow user and assistant intervals to overlap, and tune any gap timeout against recorded conversations. A brief acknowledgment from the other speaker may belong within an ongoing exchange. Grouping fragments must not trigger tool execution or cancel backend work by itself.

202 

203 

204 

205 

206 

207Keep transcript timing separate from audio playback. WebSocket `session.output_audio.delta` events have no timing fields or output-audio-done event; WebRTC delivers audio through its media track. See [Connections](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) for audio handling.

208 

209### Display captions

210 

211Build caption rows that can grow while both speakers are talking:

212 

2131. **Preserve the text.** Store each speaker's original `delta`, `start_ms`, and `end_ms`. Concatenate text exactly as received, including spaces and repeated words. Don't trim fragments or insert spaces between them.

2142. **Update each speaker independently.** Allow user and assistant rows to grow during overlapping speech. Keep earlier assistant text visible after an interruption, and start a new row when the assistant resumes.

2153. **Keep rows stable.** Assign display IDs in your application and preserve row order as text grows. Don't derive row identity from changing text or end timestamps, or move a row to the bottom whenever it receives a fragment.

2164. **Revisit grouping for late fragments.** Use transcript timestamps to group nearby fragments from the same speaker. Allow late text to update earlier rows and revise fragment assignments while retaining the original fragments. These display groups are not complete semantic turns; any gap threshold is an application choice to test.

2175. **Let the reader control scrolling.** Follow new text while the reader is at the bottom. Pause automatic scrolling when they scroll up, and provide a way to return to the latest captions.

2186. **Show tool progress in a status area.** Use assistant transcript events for spoken captions. Display tool activity and backend results outside the captions; receiving a result does not mean the assistant has said it.

219 

220Test the display with overlapping speech, short acknowledgments, interruptions, long pauses, and translation where the two speakers' text arrives at different rates.

221 

222### Control microphone input

223 

224Send `session.input_audio.mute` to mute input without ending the session:

225 

226```javascript

227/**

228 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

229 */

230export function sendUpdate(connection) {

231 connection.send({

232 type: "session.input_audio.mute",

233 event_id: "mute_1",

234 });

235}

236```

237 

238```python

239from openai.resources.live.live import AsyncLiveConnection

240from openai.resources.live.sideband import AsyncSidebandConnection

241 

242 

243async def send_update(

244 connection: AsyncLiveConnection | AsyncSidebandConnection,

245) -> None:

246 await connection.session.input_audio.mute(

247 event_id="mute_1",

248 )

249```

250 

251 

252Wait for `session.input_audio.muted` with `client_event_id: "mute_1"` before treating the command as accepted. To resume input, send `session.input_audio.unmute` and wait for `session.input_audio.unmuted`. Handle errors for either command.

253 

254Muting input does not stop inference, delegated work, or generated speech. Control microphone capture and audio playback separately in your application when those controls are needed.

255 

256### Greet before the caller speaks

257 

258To request a greeting after `session.started`:

259 

2601. Send one fresh `session.instructions.append` with `delegation_id: null`. Include the greeting, its language, and an explicit instruction to greet immediately without waiting for the caller, then pause and listen. Keep the existing startup instructions.

2612. Wait for `session.instructions.appended`, matching its `client_event_id` to your command. Handle a rejected command before continuing.

2623. Keep input audio running, including silence before the caller speaks. On WebSocket, continue sending `session.input_audio.append`; on WebRTC, keep the negotiated input audio track active. Observe output transcript and audio for the greeting.

263 

264Use the language specified by your application until the caller speaks; don't infer it from a name, phone number, or location. See [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for prompt design.

265 

266For a greeting that needs to follow application instructions, send those instructions with `session.instructions.append`, then use a short `session.commentary.append` to prompt the assistant to begin. For example: “Begin the conversation now, following the instructions provided.” Keep input audio running, including silence before the caller speaks.

267 

268Instructions request a greeting; they do not guarantee exact wording or uninterrupted playback. The API does not emit an opening-completed event, and acknowledgment does not mean the greeting was heard. Use application-controlled playback if the audio must be verbatim. Test your greeting with the languages and interruptions your application supports.

269 

270### Deliver a disclosure

271 

272Use `session.instructions.append` to request specific spoken wording for a disclosure. `session.commentary.append` may paraphrase the text. After `session.started`, for example, send:

273 

274```javascript

275/**

276 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

277 */

278export function sendUpdate(connection) {

279 connection.send({

280 type: "session.instructions.append",

281 event_id: "disclosure_1",

282 delegation_id: null,

283 content:

284 "Immediately say the following disclosure exactly and in full before responding to the caller: This call may be recorded for quality and training purposes.",

285 });

286}

287```

288 

289```python

290from openai.resources.live.live import AsyncLiveConnection

291from openai.resources.live.sideband import AsyncSidebandConnection

292 

293 

294async def send_update(

295 connection: AsyncLiveConnection | AsyncSidebandConnection,

296) -> None:

297 await connection.session.instructions.append(

298 event_id="disclosure_1",

299 delegation_id=None,

300 content=(

301 "Immediately say the following disclosure exactly and in full before "

302 "responding to the caller: This call may be recorded for quality and "

303 "training purposes."

304 ),

305 )

306```

307 

308 

309Keep input audio running as described in [Greet before the caller speaks](#greet-before-the-caller-speaks). Choose the delivery point deliberately: an instruction sent during the conversation can interrupt speech in progress.

310 

311This requests the wording; it does not guarantee exact delivery. Verify the complete spoken disclosure and actual playback before marking it delivered. `session.instructions.appended` confirms only that the instruction was accepted. If exact audio delivery is required, play a verified recording or rendered clip through your application and control GPT-Live output while it plays. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).

312 

313 

314 

315 

316 

317## Manage longer conversations

318 

319GPT-Live automatically manages context during long conversations; no configuration parameter is needed. The instructions you provide at session start are preserved throughout compaction. You don’t need to resend them.

320 

321The default context window holds 128,000 tokens, including your instructions, conversation text, and audio tokens that don’t appear in the transcript.

322 

323GPT-Live summarizes older conversation history in the background. When context usage exceeds 90%, it starts a replacement voice engine within the same session. The replacement receives your original instructions and up to 8,192 tokens of conversation history, containing recent messages and, when available, a summary of older messages. Preparing a summary does not immediately change the running engine’s context.

324 

325 

326 

327 

328 

329Older conversation details may be summarized or omitted. Keep important facts, confirmed actions, and current task state in your application, and provide relevant context when needed.

330 

331## Store and fork a session

332 

333Set `store` to `true` in the session configuration at creation to save a recording for later download or forking. Storage defaults to `false` and must be enabled for your project. Downloads and forks require a completed stored recording and a data policy that permits persistence. Recordings expire after 30 days. With Zero Data Retention, `store` is treated as `false`, and recording downloads and forks are unavailable. See [GPT-Live data controls](https://developers.openai.com/api/docs/guides/your-data#v1livesessions).

334 

335For example, add this field to the `session` object in your WebSocket `session.start` event or WebRTC creation request:

336 

337```json

338{

339 "store": true

340}

341```

342 

343Save the source session ID from `session.started` or the WebRTC creation response. A fork starts a **new session with a new ID** from the stored session state. It does not reopen the original connection or reuse the source session ID.

344 

345Start the fork through the transport your application uses:

346 

347| Transport | Start the fork |

348| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |

349| WebSocket | Connect to `wss://api.openai.com/v1/live/sessions/{source_session_id}/fork`. |

350| WebRTC | Send a new SDP offer to `POST /v1/live/sessions/{source_session_id}/fork`. Apply the returned `transport.sdp` answer to the new peer connection. |

351 

352A fork inherits the stored session configuration, subject to the transport rules below. For a WebSocket fork, send `session.start` with a required `session` object; `{}` supplies no overrides. Do not supply a new model or repeat the original instructions or input. You can override `store`, Responses delegation settings, and the new WebSocket audio format. WebRTC forks can override `store`, Responses delegation settings, and frontend client permissions. Omitting `store` on a fork inherits the source session's setting.

353 

354A WebSocket fork does **not** inherit the source audio format: set `audio.format` explicitly or use the default PCM16 at 24 kHz. It also discards inherited frontend data-channel permissions. WebRTC forks negotiate their audio format and reject `audio.format`; they preserve frontend permission settings unless you override them.

355 

356Wait for `session.started` before sending further WebSocket commands. WebRTC starts through the HTTP request and must not receive a second `session.start` on its data channel.

357 

358### Start a WebSocket fork

359 

360Set `OPENAI_API_KEY` and `OPENAI_LIVE_SESSION_ID` to your API key and stored source session ID. These examples confirm startup and then close the fork. To continue the conversation, send and receive audio after `session.started` using the [WebSocket connection flow](https://developers.openai.com/api/docs/guides/voice-websockets?api=live). See the [fork WebSocket reference](https://developers.openai.com/api/reference/resources/live/fork-websocket) for the startup fields and events.

361 

362```javascript

363import OpenAI from "openai";

364import { ForksWS } from "openai/resources/live/forks/ws";

365 

366const sourceSessionId = process.env.OPENAI_LIVE_SESSION_ID;

367if (!sourceSessionId) throw new Error("Set OPENAI_LIVE_SESSION_ID");

368 

369const ws = new ForksWS(new OpenAI(), { session_id: sourceSessionId });

370let finalized = false;

371try {

372 for await (const event of ws) {

373 if (event.type === "open") {

374 ws.send({ type: "session.start", session: {} });

375 } else if (event.type === "error") {

376 throw event.error;

377 } else if (event.type === "message") {

378 if (event.message.type === "session.started") {

379 console.log("Fork ready:", event.message.session.id);

380 // This startup example closes the fork after confirming it is ready.

381 ws.send({ type: "session.close" });

382 } else if (event.message.type === "session.closed") {

383 console.log("Final usage:", event.message.usage);

384 finalized = true;

385 break;

386 }

387 }

388 }

389 if (!finalized) throw new Error("Connection closed before session.closed");

390} finally {

391 ws.close();

392}

393```

394 

395```python

396import os

397 

398from openai import OpenAI

399 

400client = OpenAI()

401source_session_id = os.environ["OPENAI_LIVE_SESSION_ID"]

402 

403with client.live.forks.connect(session_id=source_session_id) as connection:

404 connection.session.start(session={})

405 finalized = False

406 for event in connection:

407 if event.type == "session.started":

408 print("Fork ready:", event.session.id)

409 # This startup example closes the fork after confirming it is ready.

410 connection.session.close()

411 elif event.type == "session.closed":

412 print("Final usage:", event.usage)

413 finalized = True

414 break

415 elif event.type == "error":

416 raise RuntimeError(event.error.message)

417 if not finalized:

418 raise RuntimeError("Connection closed before session.closed")

419```

420 

421 

422### Start a WebRTC fork

423 

424Create a new SDP offer in your frontend and send it to your backend. The following backend examples read that offer from the file named by `OPENAI_LIVE_SDP_OFFER_FILE` and fork the stored `OPENAI_LIVE_SESSION_ID`:

425 

426```javascript

427import OpenAI from "openai";

428import { readFile } from "node:fs/promises";

429 

430const sourceSessionId = process.env.OPENAI_LIVE_SESSION_ID;

431const offerFile = process.env.OPENAI_LIVE_SDP_OFFER_FILE;

432if (!sourceSessionId || !offerFile) {

433 throw new Error("Set OPENAI_LIVE_SESSION_ID and OPENAI_LIVE_SDP_OFFER_FILE");

434}

435const offerSdp = await readFile(offerFile, "utf8");

436 

437const client = new OpenAI();

438const fork = await client.live.sessions.fork(sourceSessionId, {

439 transport: { type: "webrtc", sdp: offerSdp },

440});

441console.log(JSON.stringify(fork));

442```

443 

444```python

445import os

446from pathlib import Path

447 

448from openai import OpenAI

449 

450client = OpenAI()

451source_session_id = os.environ["OPENAI_LIVE_SESSION_ID"]

452offer_sdp = Path(os.environ["OPENAI_LIVE_SDP_OFFER_FILE"]).read_bytes().decode()

453 

454fork = client.live.sessions.fork(

455 source_session_id,

456 transport={"type": "webrtc", "sdp": offer_sdp},

457)

458print(fork.model_dump_json())

459```

460 

461 

462Return the response to your frontend, apply `transport.sdp` as the new peer connection's answer, and retain the new `session.id`. Keep the API key on your backend.

463 

464Use the new session ID for later sideband connections and session controls. Keep application task state separately: restoring conversation state does not confirm that a pending backend action completed. Reconcile uncertain results before retrying an action. If you don't have a stored session to fork, [seed a new session with saved history](#seed-a-session-with-prior-conversation).

465 

466### Download a recording

467 

468After the stored recording is finalized, download its audio with `GET /v1/live/sessions/{session_id}/content`. The response is binary stereo WAV, with input audio in the left channel and output audio in the right channel. Set `OPENAI_LIVE_SESSION_ID` to the stored session ID. These examples stream the response to `recording.wav`:

469 

470```javascript

471import OpenAI from "openai";

472import { createWriteStream } from "node:fs";

473import { pipeline } from "node:stream/promises";

474 

475const sessionId = process.env.OPENAI_LIVE_SESSION_ID;

476if (!sessionId) throw new Error("Set OPENAI_LIVE_SESSION_ID");

477 

478const client = new OpenAI();

479const response = await client.live.sessions.downloadRecording(sessionId);

480if (!response.body) throw new Error("Recording response has no body");

481await pipeline(response.body, createWriteStream("recording.wav"));

482```

483 

484```python

485import os

486 

487from openai import OpenAI

488 

489client = OpenAI()

490session_id = os.environ["OPENAI_LIVE_SESSION_ID"]

491 

492with client.live.sessions.with_streaming_response.download_recording(

493 session_id

494) as response:

495 response.stream_to_file("recording.wav")

496```

497 

498 

499## Handle errors and end the session

500 

501Keep reading session events until the session finalizes. Distinguish a rejected command, a failed connection, and a completed session so your application can recover appropriately.

502 

503### Handle rejected commands

504 

505Read `error` events alongside acknowledgments. When present, `error.client_event_id` identifies the outgoing command that failed:

506 

507```json

508{

509 "type": "error",

510 "event_id": "event_error",

511 "error": {

512 "type": "invalid_request_error",

513 "code": "immutable_field_update",

514 "message": "The delegation type cannot change after session startup.",

515 "param": "session.delegation.type",

516 "client_event_id": "event_update"

517 }

518}

519```

520 

521An error code can be `null`, and an error may lack a client event ID. Handle those cases without assuming a command succeeded. For an immutable-field error, keep the current configuration or create a new session with the intended settings.

522 

523### Handle moderation

524 

525Moderation can affect the session in two ways:

526 

527- Some moderation events end the session.

528- Others cut off assistant audio for the remainder of its current speech and emit an `error` event without ending the session.

529 

530Read `error` events even while audio is playing. Don't assume every moderation error closes the session, or that an audio interruption means the connection failed. Keep application state aligned with the session lifecycle, and don't mark an interrupted spoken message as fully delivered. Application-level [conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) remain separate from this built-in moderation behavior.

531 

532### Usage and graceful close

533 

534`session.usage.updated` reports cumulative voice duration in seconds:

535 

536```json

537{

538 "type": "session.usage.updated",

539 "event_id": "event_usage_1",

540 "usage": { "seconds": 12 },

541 "context_window": { "usage_ratio": 0.42 }

542}

543```

544 

545These are snapshots, not increments to sum. Backend token usage is separate; preserve it from nested Responses completion events. See [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live) for usage accounting.

546 

547To close gracefully:

548 

5491. Finish any delegated Responses work your application needs, including pending function results and response continuations.

5502. Install the `session.closed` listener before sending `session.close`.

5513. Send `session.close` and stop submitting new work to the session. Keep the WebSocket or WebRTC connection, data channel, and any attached sideband receiver alive while pending session events drain.

5524. Read the final `usage.seconds`, `reason`, and session snapshot from `session.closed`. Preserve delegated usage already received through `response.event`.

5535. Clean up transports and audio devices after that event. If finalization fails or exceeds a timeout your application sets, report incomplete finalization and release the resources.

554 

555Sending `session.close` cancels queued Responses and rejects further commands. An active response can finish, but one waiting for a function result cannot continue after closing starts. Decide separately whether to finish or cancel work your application runs through client delegation.

556 

557The `session.closed` event establishes finalization; the embedded session is a configuration snapshot. A socket close alone does not establish success, and a transport close code after a valid final event does not invalidate finalization. Closing WebRTC immediately after sending the command can prevent delivery of the final event.

558 

559The final event's `reason` explains why the session ended:

560 

561| Reason | Meaning |

562| ----------------- | -------------------------------------------------------------------- |

563| `close_requested` | Your application sent `session.close` or called the hangup endpoint. |

564| `expired` | The session reached its duration limit. |

565| `content` | A safety filter ended the session. |

566| `remote_hangup` | The remote primary connection ended gracefully. |

567| `connection_lost` | The primary or upstream connection was lost unexpectedly. |

568 

569A `session.closed` event confirms finalization even when the reason is a connection loss or safety termination. Without that event, final usage remains unconfirmed. A stored session can take longer to finalize while its recording is saved; choose an application timeout that accounts for storage.

570 

571### Recover from a failed connection

572 

573An HTTP session-creation error means the session did not reach `session.started`. Handle startup errors separately from errors in a running session. If a running connection fails before `session.closed`, retain the latest observed usage and mark final usage as unconfirmed.

574 

575If a stored session is available, [fork it](#store-and-fork-a-session) to start a new session from its saved state. Otherwise, create a replacement session with relevant saved history. Reconcile pending actions with your backend before retrying them, and suppress stale results from the previous session. Restore application state explicitly rather than assuming a new connection resumes the previous session or its pending work.

guides/live-delegation.md +619 −0 created

Details

1# Delegation and tools in GPT-Live

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5GPT-Live delegates reasoning and tool use to a backend while it manages the spoken conversation. Backend work can run through the configured Responses model or, with client delegation, any model, agent, or service your application operates. In either mode, your application owns permissions, confirmations, business records, and task state.

6 

7Read more about [steering the live model for delegation and tools](https://developers.openai.com/api/docs/guides/live-prompting#delegation) in the prompting guide.

8 

9 

10 

11 

12 

13## Choose a delegation mode

14 

15With **[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#configure-responses-delegation)**, GPT-Live calls the Responses model you choose, supplies conversation context, and returns backend results to the live conversation. With **[client delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#configure-client-delegation)**, your application prepares the context, runs an agent or workflow, and sends results back to GPT-Live.

16 

17Start with Responses delegation when its managed workflow fits. Choose client delegation when you need more control over backend context, execution, or the results returned to GPT-Live.

18 

19 

20 

21 

22| Consideration | Favor Responses delegation when… | Favor client delegation when… |

23| ----------------------------- | ---------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |

24| **Implementation effort** | You want GPT-Live to prepare backend requests, manage connections, and return results to the conversation. | You want to build and operate those pieces yourself. |

25| **Reviewing backend results** | Backend output can return directly to GPT-Live. | Your application must validate, redact, combine, or discard results before they reach GPT-Live. |

26| **Backend capabilities** | Your workflow fits the Responses settings and tools supported by GPT-Live. | You need another backend, multiple models, or API capabilities beyond the managed configuration. |

27| **Context ownership** | The conversation context supplied by GPT-Live fits your application. | You need to choose exactly which history, memory, and application state each backend request receives. |

28| **Execution policy** | A configured model and tool loop fits the task. | You need custom routing between code and models, fallbacks, checkpoints, or budgets across backend steps. |

29 

30 

31 

32 

33For example, a travel assistant can send flight-status questions to an airline service and itinerary changes to a separate planning agent. The application chooses which backend to call and what verified result to return to GPT-Live.

34 

35In both modes, your application manages task state and enforces permissions and required confirmations before running its custom tools. Reviewing backend results is a separate decision: it does not approve every word GPT-Live speaks or guarantee silence while validation runs. See [Control playback when needed](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed).

36 

37Client delegation also requires your application to [maintain conversation context](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#keep-the-conversation-context-in-your-application). The delegation event contains metadata, not task text; use transcript events and application state to prepare the backend request.

38 

39Compare latency, task success, and cost on your own workload when [evaluating your voice agent](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation). For guidance specific to your existing architecture, see [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration#choose-your-delegation-mode).

40 

41Choose the mode when you create the session; to change modes, start a new session.

42 

43{/* prettier-ignore */}

44 

45 

46## Configure Responses delegation

47 

48Add this delegation configuration when [creating your Live session](https://developers.openai.com/api/docs/guides/live). Choose the Responses model independently of the voice model:

49 

50```javascript

51/** @type {import("openai/resources/live/live").SessionConfig} */

52```

53 

54```python

55from openai.types.live.session_config_param import SessionConfigParam

56 

57session: SessionConfigParam = {

58 "model": "gpt-live-1",

59 "delegation": {

60 "type": "responses",

61 "responses": {

62 "model": "gpt-5.6-terra",

63 "instructions": "[Your backend prompt]",

64 },

65 },

66}

67```

68 

69 

70Start with [GPT-5.6 Terra](https://developers.openai.com/api/docs/models/gpt-5.6-terra), or try [GPT-5.6 Luna](https://developers.openai.com/api/docs/models/gpt-5.6-luna) for cost-sensitive workloads. Compare answer quality and latency on your tasks before choosing a backend model.

71 

72Register supported tools in `delegation.responses.tools`. Use `delegation.responses.tool_choice` to control which tools the backend can use: `"auto"` lets it choose, `"required"` requires a tool call, and `"none"` prevents one. You can also select a named function. Set `delegation.responses.parallel_tool_calls` to `true` to allow independent lookups together, or `false` when calls must run sequentially. Your application still executes its custom functions and enforces dependencies and approvals. These settings do not force the live model to delegate.

73 

74The Responses configuration requires a backend `model` at creation. It supports `function` definitions and `web_search` entries in `tools`. It also exposes `max_output_tokens` (at least 16 when set), `service_tier`, and the `reasoning` and `text` settings supported by the selected backend model. See [Reduce backend latency](#reduce-backend-latency) for settings you can tune.

75 

76If [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode) is available for your model and project, consider it for latency-sensitive calls. For GPT-Live, select it with `delegation.responses.service_tier: "priority"`.

77 

78As the conversation changes, send `session.update` with changes in `session.delegation.responses` to update the backend model, instructions, available tools, `tool_choice`, or other supported settings without starting a new Live session. Omitted settings retain their values. Setting `delegation` to `null` selects client mode and cannot reset a running Responses session; switching modes fails with `immutable_field_update`.

79 

80These settings use familiar Responses concepts, but Live supports a subset of the standalone Responses API. Live supplies conversation context and initiates delegated work. Configure the backend through the session; the Live `response.create` command uses that configuration and does not accept a standalone Responses request body.

81 

82## Steer the live conversation from your application

83 

84Responses delegation manages the backend workflow, but your application can still send context directly to the GPT-Live model. If you monitor the call through a sideband WebSocket or the main event connection, you can use `session.instructions.append`, `session.thinking.append`, or `session.commentary.append` with `delegation_id: null`. For example, a transcript-based guardrail can append an instruction to redirect the conversation. This steers the live model; it does not change the Responses backend prompt or cancel work already in progress.

85 

86## Handle Responses delegation

87 

88For Responses-backed work, `session.delegation.created` has `target: "responses"` and a `response_id`. Subsequent Responses events arrive inside a `response.event` envelope:

89 

90```json

91{

92 "type": "response.event",

93 "event_id": "event_response_1",

94 "delegation_id": "item_9tA2cB6n2V8c4X1z7Q5r9",

95 "event": {

96 "type": "response.output_text.delta",

97 "sequence_number": 4,

98 "item_id": "msg_123",

99 "output_index": 0,

100 "content_index": 0,

101 "delta": "The forecast is",

102 "logprobs": []

103 }

104}

105```

106 

107Dispatch on `envelope.event.type` and preserve the outer `delegation_id`. Do not handle every top-level `response.*` value as an unwrapped Responses event. Tolerate additional nested Responses lifecycle events.

108 

109Live speech and delegated work continue independently. A completed backend response does not itself mean the user heard the answer. Use the Live output transcript and audio for the spoken part of the interaction.

110 

111### Complete a client-actionable function call

112 

113Read completed function calls from nested `response.output_item.done` events. The finished function item contains `call_id`, `name`, and `arguments`; an arguments-done event alone is not sufficient to identify the call.

114 

115Track the response ID from nested `response.created` alongside the outer `delegation_id`, and collect that response's function calls from `response.output_item.done`. Forwarded lifecycle snapshots deliberately contain `response.output: []`, including at `response.completed`; their `tools` array is empty, `instructions` is `null`, and `input` is omitted. An empty terminal output list does **not** mean there are no pending function calls. Use the collected calls to determine which results must be submitted before continuing.

116 

117After executing the authorized operation, append the result as a Responses item:

118 

119```javascript

120/**

121 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

122 */

123export function sendUpdate(connection) {

124 connection.send({

125 type: "response.item.create",

126 event_id: "tool_result_1",

127 item: {

128 type: "function_call_output",

129 call_id: "call_123",

130 output: '{"status":"confirmed","order_id":"order_123"}',

131 },

132 });

133}

134```

135 

136```python

137from openai.resources.live.live import AsyncLiveConnection

138from openai.resources.live.sideband import AsyncSidebandConnection

139from openai.types.responses.response_input_item_param import ResponseInputItemParam

140 

141 

142async def send_update(

143 connection: AsyncLiveConnection | AsyncSidebandConnection,

144) -> None:

145 item: ResponseInputItemParam = {

146 "type": "function_call_output",

147 "call_id": "call_123",

148 "output": '{"status":"confirmed","order_id":"order_123"}',

149 }

150 await connection.response.item.create(

151 event_id="tool_result_1",

152 item=item,

153 )

154```

155 

156 

157Then explicitly continue the response:

158 

159```javascript

160/**

161 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

162 */

163export function sendUpdate(connection) {

164 connection.send({

165 type: "response.create",

166 event_id: "continue_1",

167 });

168}

169```

170 

171```python

172from openai.resources.live.live import AsyncLiveConnection

173from openai.resources.live.sideband import AsyncSidebandConnection

174 

175 

176async def send_update(

177 connection: AsyncLiveConnection | AsyncSidebandConnection,

178) -> None:

179 await connection.response.create(

180 event_id="continue_1",

181 )

182```

183 

184 

185Submit every required result for the pending tool calls before continuing. Appending a function result does not automatically continue the response. `response.item.create` has no standalone success acknowledgment; keep processing errors and the subsequent nested response lifecycle.

186 

187`response.create` is a Live command for creating or continuing delegated Responses work, using the session's configured backend. Do not attach a Responses API creation body, backend model override, or `delegation_id` to this event. Both commands require Responses delegation.

188 

189

190 

191

192 

193 

194## Configure client delegation

195 

196Set `delegation` when [creating your Live session](https://developers.openai.com/api/docs/guides/live):

197 

198```javascript

199/** @type {import("openai/resources/live/live").SessionConfig} */

200```

201 

202```python

203from openai.types.live.session_config_param import SessionConfigParam

204 

205session: SessionConfigParam = {"model": "gpt-live-1", "delegation": {"type": "client"}}

206```

207 

208 

209This selects client delegation for the session. Configure the backend separately: your application chooses its model or service, instructions, tools, and how to route work. If you use the Responses API for that backend, set its model and tools in your own Responses requests. The Live session does not configure or run those backend tools.

210 

211When GPT-Live requests help, your application builds the backend request from conversation and application context, runs the work, and decides which results to send back. Enforce permissions and required confirmations before executing your tools. Retain the full conversation history in your application so you can provide the relevant context for each backend request.

212 

213## Keep the conversation context in your application

214 

215For client delegation, **collect transcripts and keep the current task state yourself**.

216 

217Listen for `session.input_transcript.delta` and `session.output_transcript.delta`. These events contain transcript text in `delta`, along with `start_ms` and `end_ms` timestamps. Keep enough history to understand short replies such as “yes,” corrections such as “Thursday, not Friday,” and details supplied earlier. A transcript fragment is not a complete user turn, and transcripts may contain mistakes.

218 

219The separate `session.delegation.created` event contains an `offset_ms` timestamp and delegation metadata, including `delegation.id` and `delegation.target`. It does **not** contain the user's utterance or task text. Use the transcript events and application state to work out what the user wants. Save `delegation.id` so you can match updates to that request.

220 

221Keep long records and full tool output in the backend. If you create a replacement session, restore the relevant context from your application and check which actions already ran before repeating any work.

222 

223### Receive a client delegation

224 

225`session.delegation.created` identifies a delegation:

226 

227```json

228{

229 "type": "session.delegation.created",

230 "event_id": "event_delegation",

231 "offset_ms": 1000,

232 "delegation": {

233 "id": "item_9tA2bF3h7K9m2P5q8R1s4",

234 "type": "delegation",

235 "target": "client"

236 }

237}

238```

239 

240Read `event.delegation.id`. The delegation object contains metadata, not task text. Maintain the transcript and application context needed by your own delegated-work handler. Current IDs have an `item_` prefix, as illustrated here; treat the full ID as opaque and return it unchanged rather than constructing or parsing one.

241 

242Return a result using that ID:

243 

244```javascript

245/**

246 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

247 */

248export function sendUpdate(connection) {

249 connection.send({

250 type: "session.commentary.append",

251 event_id: "result_123",

252 delegation_id: "item_9tA2bF3h7K9m2P5q8R1s4",

253 content: "The order shipped today and should arrive tomorrow.",

254 });

255}

256```

257 

258```python

259from openai.resources.live.live import AsyncLiveConnection

260from openai.resources.live.sideband import AsyncSidebandConnection

261 

262 

263async def send_update(

264 connection: AsyncLiveConnection | AsyncSidebandConnection,

265) -> None:

266 await connection.session.commentary.append(

267 event_id="result_123",

268 delegation_id="item_9tA2bF3h7K9m2P5q8R1s4",

269 content="The order shipped today and should arrive tomorrow.",

270 )

271```

272 

273 

274Use `session.thinking.append` to add information to the model's internal reasoning without speaking it aloud when appended. Use `session.commentary.append` for a result the model should speak aloud; the model is trained to paraphrase the appended text. All appends contain a plain string and require `delegation_id`, including when its value is `null`. A non-null ID must name a known client delegation.

275 

276Repeated result appends can continue the same client delegation. An appended acknowledgment arrives after estimated context injection; it is not proof that the model has consumed or spoken the result, or that an external action succeeded.

277 

278 

279 

280## Start with your existing backend prompt

281 

282Use your existing text-agent prompt as a starting point. Keep its task instructions and business rules with the backend, and adapt instructions that assume a text chat or direct control of speech. Explain how to handle voice transcripts and return useful results. Enforce permissions and required confirmations in your application.

283 

284```text

285## Voice conversation context

286You are helping an assistant in a live voice conversation. Transcripts

287can contain mistakes, unfinished phrases, and later corrections. Use

288the latest context and verified records. If a needed detail is still

289unclear, ask for that detail instead of guessing.

290 

291## Task instructions

292[Your task instructions, business rules, available tools,

293and confirmation requirements.]

294 

295## Return the result

296Return the relevant facts, whether the task is complete, and what comes next.

297Use confirmed values. Do not invent a successful action.

298```

299 

300Keep large structured payloads, lengthy tool output, and Markdown intended for display in the backend. Give GPT-Live the relevant facts and let it choose how to say them. A concise tool result doesn't need an additional model call to rewrite it for speech.

301 

302With client delegation, [return the result directly to GPT-Live](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#receive-a-client-delegation). With Responses delegation, follow the [function-result flow](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#complete-a-client-actionable-function-call) to continue backend work.

303 

304SDK event examples below use `connection`, a connected primary Live WebSocket or sideband from the [connection guides](https://developers.openai.com/api/docs/guides/voice-websockets?api=live). Call the helper after `session.started` on a primary connection; an attached sideband already belongs to a running session.

305 

306## Send the right kind of update

307 

308Choose an event based on how GPT-Live should use the content:

309 

310| What you want to send | Event |

311| ----------------------------------------------------------------------------------------------------------- | ----------------------------- |

312| System-level instructions for the live model, such as a greeting, disclosure, or direction to stop speaking | `session.instructions.append` |

313| Information for internal reasoning, not spoken on append but usable for relevant user questions | `session.thinking.append` |

314| Information the model should speak aloud, paraphrasing the appended text | `session.commentary.append` |

315 

316All three use a plain-string `content`, limited to 500 tokens per append. Include `delegation_id`: use the original client delegation ID for an update about that task, or `null` for general session context. A non-null ID must identify a known client delegation. Instructions still apply to the live session; an ID does not turn them into a separate backend prompt.

317 

318An appended instruction can interrupt the model's current speech or behavior. Use it when the application needs to redirect the conversation; enforce any related tool or action block in application state.

319 

320For quiet progress during a client-managed task:

321 

322```javascript

323/**

324 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

325 */

326export function sendUpdate(connection) {

327 connection.send({

328 type: "session.thinking.append",

329 event_id: "availability_progress",

330 delegation_id: "item_123",

331 content: "Checking Thursday availability. No appointment has been booked.",

332 });

333}

334```

335 

336```python

337from openai.resources.live.live import AsyncLiveConnection

338from openai.resources.live.sideband import AsyncSidebandConnection

339 

340 

341async def send_update(

342 connection: AsyncLiveConnection | AsyncSidebandConnection,

343) -> None:

344 await connection.session.thinking.append(

345 event_id="availability_progress",

346 delegation_id="item_123",

347 content="Checking Thursday availability. No appointment has been booked.",

348 )

349```

350 

351 

352For a confirmed booking, send the result the user should hear:

353 

354```javascript

355/**

356 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

357 */

358export function sendUpdate(connection) {

359 connection.send({

360 type: "session.commentary.append",

361 event_id: "appointment_result",

362 delegation_id: "item_123",

363 content: "Your appointment is confirmed for Thursday at 2:00 PM",

364 });

365}

366```

367 

368```python

369from openai.resources.live.live import AsyncLiveConnection

370from openai.resources.live.sideband import AsyncSidebandConnection

371 

372 

373async def send_update(

374 connection: AsyncLiveConnection | AsyncSidebandConnection,

375) -> None:

376 await connection.session.commentary.append(

377 event_id="appointment_result",

378 delegation_id="item_123",

379 content="Your appointment is confirmed for Thursday at 2:00 PM",

380 )

381```

382 

383 

384Only send that result after the booking has actually succeeded. For a session-wide instruction, use `session.instructions.append` with `delegation_id: null`.

385 

386For example, after your application blocks a request under its guardrails, you can redirect the conversation:

387 

388```javascript

389/**

390 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

391 */

392export function sendUpdate(connection) {

393 connection.send({

394 type: "session.instructions.append",

395 event_id: "guardrail_block_17",

396 delegation_id: null,

397 content:

398 "Stop speaking about that request. Briefly explain that you cannot help with it, then wait for the user.",

399 });

400}

401```

402 

403```python

404from openai.resources.live.live import AsyncLiveConnection

405from openai.resources.live.sideband import AsyncSidebandConnection

406 

407 

408async def send_update(

409 connection: AsyncLiveConnection | AsyncSidebandConnection,

410) -> None:

411 await connection.session.instructions.append(

412 event_id="guardrail_block_17",

413 delegation_id=None,

414 content=(

415 "Stop speaking about that request. Briefly explain that you cannot help "

416 "with it, then wait for the user."

417 ),

418 )

419```

420 

421 

422The instruction does not cancel backend work. [Block the affected action and handle any work already running](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) in your application.

423 

424The corresponding acknowledgements are `session.thinking.appended`, `session.commentary.appended`, and `session.instructions.appended`. Match their `client_event_id` to your outgoing `event_id`. The acknowledgment waits for estimated context injection, not for speech or playback to finish. See [when context reaches the model](https://developers.openai.com/api/docs/guides/live-conversations#understand-when-context-reaches-the-model) for timing and error handling.

425 

426Quiet context can still affect what the model says later. It is not a private place for secrets or hidden reasoning. Send useful facts and brief progress summaries.

427 

428## Keep updates accurate and useful

429 

430During longer tasks, send an update when something useful changes: a step finishes, a delay matters, or the user needs to answer a question.

431 

432Use `session.thinking.append` for background progress in client mode. Use `session.commentary.append` when the update is useful to say aloud.

433 

434For spoken updates, send `session.commentary.append` with content that matches the task's verified state:

435 

436| State | Example content |

437| ---------------------- | ------------------------------------------ |

438| Still working | “I'm checking the available appointments.” |

439| Completed | “You're booked for Thursday at 2:00 PM.” |

440| Failed | “That time is no longer available.” |

441| Cancellation confirmed | “Your appointment has been canceled.” |

442 

443A spoken interruption does not automatically cancel backend work. If the user changes Friday to Thursday, update the active task and ignore late Friday results. Your application must decide whether to cancel work, change it, or let it finish. Check that cancellation succeeded before saying it did.

444 

445Before retrying a failed tool call, check whether the original action already happened. For example, a lost response should not cause a second booking. If the outcome is unclear, say so and offer the next useful step.

446 

447## Share UI context

448 

449Give GPT-Live a concise summary of the current page or task, relevant selections, and facts that help interpret references such as “this option.” Build the summary directly from application state; no extra model call is needed to format it.

450 

451Send UI context at session start and when relevant state changes. Skip unchanged updates and combine rapid changes into a short summary of the latest state. Make changes to previous selections explicit:

452 

453- **Initial context:** “The user is reviewing a restaurant reservation: August 6 at 7 PM, two guests. No reservation has been made.”

454- **Correction:** “The selected time is now 8 PM; the previous selection was 7 PM.”

455 

456In either delegation mode, use `session.thinking.append` with `delegation_id: null` for [background context updates](https://developers.openai.com/api/docs/guides/live-conversations#add-context-during-the-conversation). Keep full HTML, DOM trees, large JSON payloads, and interaction logs in your application or backend. Treat page content as reference data, not instructions.

457 

458### Accept typed input

459 

460If a caller types an exact value, such as an order number, pass it to the backend that handles the task. A voice-only application does not need this path. Keep the typed value as user data rather than a live-model instruction.

461 

462{/* prettier-ignore */}

463 

464 

465With Responses delegation, queue a user message for the backend:

466 

467```javascript

468/**

469 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

470 */

471export function sendUpdate(connection) {

472 connection.send({

473 type: "response.item.create",

474 event_id: "typed_order_number",

475 item: {

476 type: "message",

477 role: "user",

478 content: [

479 {

480 type: "input_text",

481 text: "My order number is A0042.",

482 },

483 ],

484 },

485 });

486}

487```

488 

489```python

490from openai.resources.live.live import AsyncLiveConnection

491from openai.resources.live.sideband import AsyncSidebandConnection

492from openai.types.responses.response_input_item_param import ResponseInputItemParam

493 

494 

495async def send_update(

496 connection: AsyncLiveConnection | AsyncSidebandConnection,

497) -> None:

498 item: ResponseInputItemParam = {

499 "type": "message",

500 "role": "user",

501 "content": [{"type": "input_text", "text": "My order number is A0042."}],

502 }

503 await connection.response.item.create(

504 event_id="typed_order_number",

505 item=item,

506 )

507```

508 

509 

510Send `response.create` when ready to run or continue the backend. If it is waiting for function results, return all required results first. Queuing text does not itself cancel work already running.

511 

512

513 

514

515 

516 

517With client delegation, send the typed value directly to the backend that handles the conversation. If it corrects a running task, update that task instead of starting the same work again. You can mirror a short factual summary into the live session with `session.thinking.append`, or use `session.commentary.append` for a result the user should hear.

518 

519 

520 

521## Add images and visual context

522 

523To help a caller discuss a photo or screen, send the image and relevant context from your application to a vision-capable backend. The backend interprets the image and returns relevant text for GPT-Live to use in conversation. The Live audio frontend does not accept images directly.

524 

525{/* prettier-ignore */}

526 

527 

528With Responses delegation, configure a vision-capable backend model. Queue a supported Responses image input item with `response.item.create`, then send `response.create` to run or resume backend work. Return all required pending function results before continuing. See [Handle Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#handle-responses-delegation).

529 

530

531 

532

533 

534 

535With client delegation, send visual input to the backend that handles delegated requests, alongside the relevant conversation and application state. Return concise findings using the [client result flow](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=client#receive-a-client-delegation).

536 

537 

538 

539Keep backend image input separate from `session.input`, which seeds the Live frontend with text history at startup. See [Images and vision](https://developers.openai.com/api/docs/guides/images-vision) for supported image formats and model limitations.

540 

541## Reduce backend latency

542 

543Reduce the time between a request for backend work and a useful result for the conversation. Measure [latency at each stage](https://developers.openai.com/api/docs/guides/voice-agents#measure-latency) to locate delays. Compare useful spoken response time and task success on the same scenarios, and see the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) for evaluation guidance.

544 

545{/* prettier-ignore */}

546 

547 

548### Responses delegation

549 

550Live manages persistent WebSocket connections to Responses, prepares the connection and known request configuration in advance, and reuses prior response state when available. You do not need to implement those steps for the hosted backend. Reuse depends on the active connection and compatible state; it does not guarantee a cache hit or a specific latency.

551 

552Tune the backend through `delegation.responses`:

553 

554- `model`: choose the model that handles reasoning and tool selection independently of the voice model.

555- `reasoning.effort`: balance reasoning time and task quality using values supported by that model.

556- `service_tier`: use `auto`, `default`, `flex`, or `priority`, subject to model support and project access. `auto` follows the project's configuration. Evaluate the performance and cost of the tier you choose.

557 

558Update supported settings during the session with `session.update`. Your custom tools still run in your application, so slow service calls, queues, and tool-result buffering can delay the answer even when Live manages the Responses connection. Return each required tool result promptly and [continue the backend response](https://developers.openai.com/api/docs/guides/live-delegation?delegation-mode=responses#complete-a-client-actionable-function-call).

559 

560

561 

562

563 

564 

565### Client delegation

566 

567Your application owns the path from delegation receipt to returning a result. Prepare that path while the voice session runs:

568 

569- **Reuse backend connections.** Keep the API client and its connection pool alive across delegations. For repeated Responses calls, consider a persistent [Responses WebSocket](https://developers.openai.com/api/docs/guides/websocket-mode).

570- **Prepare known configuration.** Initialize instructions, tools, and connections before the first request needs them. Responses WebSocket mode also supports warming up known request state before generation; follow its [setup guidance](https://developers.openai.com/api/docs/guides/websocket-mode#connect-and-create-responses).

571- **Stream useful results.** Return coherent, verified chunks with `session.commentary.append`. Use `session.thinking.append` for quiet progress. Preserve the client delegation ID and the 500-token limit per append. Keep private reasoning in the backend and confirm actions before announcing success.

572- **Keep reusable input stable.** Preserve instructions, tool definitions and ordering, and unchanged history prefixes. Append new information after reusable content when your backend supports caching and continuation.

573- **Avoid unnecessary buffering.** Forward a useful result as soon as it is ready. Buffer only enough to classify the output and form a coherent chunk. Prefer structured phase metadata; if you use text prefixes to distinguish progress from results, wait for the complete prefix before forwarding text.

574 

575Measure the first useful spoken answer when comparing this path with Responses delegation.

576 

577 

578 

579### React to transcript fragments

580 

581Processing transcript fragments in your application is optional and works with either delegation mode. User and assistant [transcript fragments](https://developers.openai.com/api/docs/guides/live-conversations#transcript-deltas) arrive over WebSocket or the WebRTC data channel. You can process them with application logic or a lightweight model to start work before a delegation event arrives, or use the transcript itself to trigger application-owned work.

582 

583Use this pattern to:

584 

585- **Reduce waiting.** Start a speculative lookup when enough information is available—for example, checking availability while the user continues describing their preferences.

586- **Run guardrails.** Check the growing transcript for requests or responses that need intervention. See [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails).

587- **Adapt the conversation.** Look for wording that suggests confusion or frustration, then adjust the experience or send a focused instruction.

588- **Update the interface.** Highlight relevant controls, populate suggested fields, or show results as they become available.

589 

590For browser applications, use the WebRTC data channel for captions and local UI updates. When transcript processing runs on your server—for guardrails, lightweight model checks, or speculative tool calls—use a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) to receive events and steer the same GPT-Live session directly.

591 

592Process accumulated text when meaningful new information arrives. A fragment may be incomplete, and later speech can change the request. Discard outdated results, coordinate with subsequent delegated work to avoid duplicate actions, and apply your usual permission and confirmation checks before consequential actions.

593 

594To feed information back into the conversation:

595 

596| Intent | Event |

597| ------------------------------------------------------------- | ----------------------------- |

598| Change the live model's behavior or redirect the conversation | `session.instructions.append` |

599| Provide quiet context for subsequent responses | `session.thinking.append` |

600| Provide information the model should say aloud | `session.commentary.append` |

601 

602For updates outside a client delegation, use `delegation_id: null`. These appends steer the live model; your application controls UI changes, tool execution, and cancellation. See [Send the right kind of update](#send-the-right-kind-of-update) for append examples.

603 

604### Shared optimizations

605 

606Both delegation modes benefit from the same backend improvements:

607 

608- **Choose the model and reasoning effort for the task.** Compare configurations that meet your accuracy requirements. Use lower reasoning effort when it completes the task reliably.

609- **Keep answers concise.** Return the facts and status GPT-Live needs to continue the conversation. Avoid long explanations and extra model calls just to rewrite results for speech.

610- **Reduce tool delays and unnecessary calls.** Start authorized work when its inputs are ready, reuse results while they remain valid, and avoid repeating a completed lookup.

611- **Run independent work concurrently.** Independent lookup calls can run together. Respect dependencies and required confirmations for actions. `parallel_tool_calls` lets a model request multiple calls; your application still schedules and executes its custom functions.

612 

613See [Latency optimization](https://developers.openai.com/api/docs/guides/latency-optimization) for general Responses guidance and [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) for reusing stable input.

614 

615## Verify the complete interaction

616 

617Test both the authoritative application state and the audio the client played. A backend response can finish while the spoken result is interrupted, and a context acknowledgment confirms acceptance rather than playback. Keep operation IDs and task revisions separate from delegation IDs so reconnects, retries, and late results do not repeat or reverse an action.

618 

619Use [Evaluating voice agents](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) for repeatable tests. For an existing Realtime tool loop or chained backend, follow [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration).

guides/live-migration.md +400 −0 created

Details

1# Migrate to GPT-Live

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5GPT-Live handles the voice conversation while a backend handles task reasoning and tools. Keep your application logic, tool implementations, permissions, and durable state. The migration connects those responsibilities to the new voice interface.

6 

7This guide uses an appointment assistant: check availability, ask the user to confirm a slot, then book it. Start with a connected session from [Getting started](https://developers.openai.com/api/docs/guides/live), and keep representative conversations from your existing application for comparison.

8 

9## Before you migrate

10 

11Record the requirements your migrated application must preserve:

12 

13- **Tools and business rules:** List your existing prompts, tools, and workflows, including the conditions for each action.

14- **Input types:** Identify where audio, typed text, and images enter your application and which backend needs them. See [Add images and visual context](https://developers.openai.com/api/docs/guides/live-delegation#add-images-and-visual-context).

15- **Decisions that depend on audio:** Identify decisions that need the original sound, beyond the words in a transcript. See [Preserve decisions that depend on audio](https://developers.openai.com/api/docs/guides/live-migration?migration-path=realtime#preserve-decisions-that-depend-on-audio).

16- **Speech and playback:** Specify when speech may start, when it must stop, and which checks must finish before audio plays.

17- **Permissions and guardrails:** List authorization, confirmation, and input/output checks, and where your application enforces them. See [Adapt your guardrails](#adapt-your-guardrails).

18- **Durable state:** Identify the records, task progress, and pending actions your application must keep across disconnects and new sessions.

19- **Baseline conversations:** Save representative conversations and their starting state, expected tool actions, final application state, and spoken responses from your current application.

20 

21Use [Getting started](https://developers.openai.com/api/docs/guides/live) for session setup and the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) to plan your comparison.

22 

23## Choose your delegation mode

24 

25Your existing architecture is a useful starting point:

26 

27- **Responses delegation** fits a Realtime app where the model selects functions and your application executes them. A hosted Responses model takes over task reasoning and tool selection.

28- **Client delegation** fits an existing text agent or orchestrator. Your application supplies context, invokes that backend, and returns results to GPT-Live.

29 

30Either migration path can use either mode. For example, a Realtime app that already has a separate backend agent may keep it with client delegation. Also consider how much control you need over backend context, execution, and reviewing results before they reach GPT-Live. See [Choose a delegation mode](https://developers.openai.com/api/docs/guides/live-delegation#choose-a-delegation-mode) for the full comparison.

31 

32## Choose your migration path

33 

34Select the path that matches the application you have today.

35 

36 

37 

38## From Realtime API

39 

40Start with the [GPT-Live prompting guide](https://developers.openai.com/api/docs/guides/live-prompting). Split your existing prompt between the voice model and the backend instead of copying it wholesale into `session.instructions`. Keep conversation style and delegation guidance in the voice prompt; move detailed workflows and tool-use instructions to the backend.

41 

42**Before:** the Realtime model handles speech and selects functions such as `check_availability` and `book_appointment`. Your application executes the functions and returns their results.

43 

44**After:** GPT-Live handles speech and delegates task work. The backend selects the same functions; your application still validates and executes them. The steps here use Responses delegation. If you retain an external agent, use the [client adapter](https://developers.openai.com/api/docs/guides/live-migration?migration-path=text-agent#connect-your-existing-agent) instead.

45 

46### How Responses delegation works

47 

48Configure the backend model, instructions, and tools in `delegation.responses`. When GPT-Live decides a request needs backend work, the Live service calls that Responses model and supplies relevant conversation context. The backend reasons about the task and selects tools. Your application still runs custom functions, enforces permissions, and returns their results.

49 

50For the appointment assistant:

51 

521. The user asks which appointments are available on Friday, and GPT-Live delegates the request.

532. The Responses backend requests `check_availability`.

543. Your application runs the function, returns its result, and continues the backend response.

554. GPT-Live uses the answer from the backend to discuss available slots with the user.

56 

57GPT-Live can keep the conversation going while backend work runs. Finishing that work does not mean the assistant has finished speaking. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation) for configuration and the full event flow.

58 

59### Preserve decisions that depend on audio

60 

61Check whether existing tool decisions depend on acoustic evidence, such as a voicemail beep or a recorded greeting's cadence. GPT-Live hears the incoming audio, but its voice frontend delegates work instead of issuing ordinary structured function calls. In client mode, `session.delegation.created` carries metadata and timing, without raw audio, task text, or parsed tool arguments. A delegated backend does not automatically receive the waveform.

62 

63For answering-machine detection, explicitly route incoming audio to an audio-capable detector. One application-managed architecture to evaluate runs a separate Realtime session alongside GPT-Live for part of the call:

64 

651. Send a copy of the incoming call audio to both sessions.

662. Have the detector report its classification through a structured function call. Check each result against your schema, reject stale results, and keep an unknown state when evidence is insufficient. Allow later evidence to revise the decision.

673. Send relevant trusted context to GPT-Live, and apply your application's policy to outgoing audio playback.

68 

69Keep human or machine classification separate from recording readiness. Recognizing voicemail does not establish that the greeting and beep have finished or that recording can begin. A classifier result or context acknowledgment also does not establish permission to play audio. Use [Adapt your guardrails](#adapt-your-guardrails) and the [playback controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#control-playback-when-needed) to enforce that decision in the audio path your application controls.

70 

71Test a short “hello” that develops into a voicemail greeting, call-screening prompts, and a person picking up during voicemail. If you stop the detector before the call ends, test later human pickup after it stops. Choose when to stop the detector based on these tests and its added cost. The first human classification alone does not establish that further detection is unnecessary.

72 

73### Adapt the connection and audio lifecycle

74 

75Replace Realtime session setup with the [GPT-Live connection procedure](https://developers.openai.com/api/docs/guides/live). Recheck your transport's startup and audio format. WebRTC carries audio on media tracks and JSON events on the data channel. A primary WebSocket carries audio in JSON events.

76 

77If your Realtime application uses a server connection to monitor the call or enforce guardrails, adapt it to the [GPT-Live sideband connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#attach-to-the-existing-session). Follow [Adapt your guardrails](#adapt-your-guardrails) for the changes to conversation checks and playback.

78 

79| Existing Realtime behavior | GPT-Live adaptation |

80| ----------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |

81| Send WebSocket audio with `input_audio_buffer.append`. | Send `session.input_audio.append`; its `audio` field contains base64 raw audio. |

82| Play `response.output_audio.delta` from its `delta` field. | Play `session.output_audio.delta` from its `delta` field, in order. |

83| Commit audio or create a response to start a turn when using manual turn control. | Stream audio continuously. GPT-Live decides when to speak; remove manual audio commits and voice-turn triggers. |

84| Track audio generation and response completion with `response.output_audio.done` and `response.done`. | GPT-Live has no corresponding event marking the end of each spoken response. Track playback in your client. |

85| Display user captions from input transcription events. | Append `session.input_transcript.delta` text to the user's captions. |

86| Display assistant captions from `response.output_audio_transcript.delta`. | Append `session.output_transcript.delta` text to the assistant's captions. |

87 

88**Generation and playback:** In Realtime, `response.output_audio.done` marks the end of audio generation, while `response.done` marks the end of the response stream. These events can also occur when a response is interrupted or unsuccessful; check `response.status` in `response.done`. Neither confirms that buffered audio has finished playing. For example, the server can finish generating while the client still has a second of audio to play. Drive a "speaking" indicator from playback state.

89 

90**Captions:** Input transcription represents the user's speech; output transcription represents the assistant's generated speech. When input transcription is enabled, Realtime sends updates through `conversation.item.input_audio_transcription.delta` and a final transcript through `conversation.item.input_audio_transcription.completed`. A `delta` is a new text fragment. In GPT-Live, append each fragment to the corresponding speaker's captions independently because listening and speaking can overlap. A fragment is not a complete turn or confirmation of playback. See [Display captions](https://developers.openai.com/api/docs/guides/live-conversations#display-captions) for a display recipe.

91 

92In GPT-Live, `response.create` starts or continues delegated Responses work. It does not grant permission for the voice model to speak. For startup, greetings, interruptions, and closing a session, follow [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).

93 

94### Split conversation and backend instructions

95 

96Move conversation style and delegation guidance into `session.instructions`. Move business rules and tool-use instructions into `delegation.responses.instructions`. For a backend you run yourself, keep those rules in its existing prompt.

97 

98**Before: one Realtime prompt**

99 

100```text

101Help callers book appointments. Speak briefly. Check availability with the tool,

102ask the caller to confirm a slot, then book it. Never claim an unverified booking.

103```

104 

105**After: GPT-Live conversation instructions**

106 

107```text

108Help callers book appointments. Keep spoken replies brief. Delegate availability

109checks and booking requests. Ask the caller to confirm the proposed slot.

110Only announce a booking when the backend reports that it succeeded.

111```

112 

113**After: backend instructions**

114 

115```text

116Use the appointment tools to check current availability. Before booking, verify

117that the caller confirmed the exact slot and still has permission to book it.

118Apply the latest correction. Return verified availability, booking, or failure

119status with the date, time, and time zone.

120```

121 

122Enforce confirmation and permission checks in your application before executing a tool. Prompt instructions guide the models; they do not enforce those checks. See [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for prompt design.

123 

124### Adapt your function handlers

125 

126Keep the implementation of `check_availability` and `book_appointment`. Move their definitions from Realtime's `session.tools` or `response.tools` to `delegation.responses.tools`, using the Responses function schema. Move tool-selection settings to `delegation.responses.tool_choice` and `delegation.responses.parallel_tool_calls`. See [Configure Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation#configure-responses-delegation).

127 

128The function still returns a result for its original `call_id`. What changes is where your handler receives the call and sends the result:

129 

130| Step | Realtime API | GPT-Live with Responses delegation |

131| ------------------------------------ | -------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |

132| Receive the completed function call. | Read `response.output_item.done`. | Unwrap `response.event`, then read its inner `response.output_item.done`. |

133| Identify and execute the operation. | Read the item's `name`, `arguments`, and `call_id`; run your authorized handler. | Keep that handler and its checks. Preserve the outer `delegation_id` and backend response ID in your application. |

134| Return each function result. | Send `conversation.item.create`. | Send `response.item.create`. |

135| Continue after all required results. | Send `response.create`. | Send `response.create` to continue backend work. |

136 

137For example, after `check_availability` returns one verified slot, your result changes as follows. These are messages on an already-connected session; `call_availability` stands for the actual call ID you received.

138 

139**Before: Realtime result**

140 

141```json

142{

143 "type": "conversation.item.create",

144 "item": {

145 "type": "function_call_output",

146 "call_id": "call_availability",

147 "output": "{\"available\":true,\"slot_id\":\"slot_friday_14\",\"booked\":false}"

148 }

149}

150```

151 

152**After: GPT-Live result**

153 

154```javascript

155/**

156 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

157 */

158export function sendUpdate(connection) {

159 connection.send({

160 type: "response.item.create",

161 event_id: "availability_result_1",

162 item: {

163 type: "function_call_output",

164 call_id: "call_availability",

165 output: '{"available":true,"slot_id":"slot_friday_14","booked":false}',

166 },

167 });

168}

169```

170 

171```python

172from openai.resources.live.live import AsyncLiveConnection

173from openai.resources.live.sideband import AsyncSidebandConnection

174from openai.types.responses.response_input_item_param import ResponseInputItemParam

175 

176 

177async def send_update(

178 connection: AsyncLiveConnection | AsyncSidebandConnection,

179) -> None:

180 item: ResponseInputItemParam = {

181 "type": "function_call_output",

182 "call_id": "call_availability",

183 "output": '{"available":true,"slot_id":"slot_friday_14","booked":false}',

184 }

185 await connection.response.item.create(

186 event_id="availability_result_1",

187 item=item,

188 )

189```

190 

191 

192After submitting every required function result, continue the backend:

193 

194```javascript

195/**

196 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

197 */

198export function sendUpdate(connection) {

199 connection.send({

200 type: "response.create",

201 event_id: "continue_availability_1",

202 });

203}

204```

205 

206```python

207from openai.resources.live.live import AsyncLiveConnection

208from openai.resources.live.sideband import AsyncSidebandConnection

209 

210 

211async def send_update(

212 connection: AsyncLiveConnection | AsyncSidebandConnection,

213) -> None:

214 await connection.response.create(

215 event_id="continue_availability_1",

216 )

217```

218 

219 

220For the initial migration, setting `parallel_tool_calls` to `false` simplifies result handling. Collect calls from completed output-item events even if a terminal lifecycle snapshot has `output: []`. An arguments-done event alone does not supply the function name and `call_id`. Follow the complete [function-result procedure](https://developers.openai.com/api/docs/guides/live-delegation#complete-a-client-actionable-function-call) for collection, output submission, and errors.

221 

222### Preserve context and apply corrections

223 

224Responses delegation supplies relevant voice conversation context to the backend. Keep the authoritative appointment state in your application: selected slot, confirmed slot, permissions, active operation, and outcome. Live conversation history can be compacted; it is not your booking record.

225 

226If the user says “Actually, Friday instead” while a Thursday lookup is pending, update the task's revision and invalidate the earlier slot confirmation. Before executing a booking, check that its arguments still match the current task and confirmation. Return an accurate superseded or cancelled result for any pending function call your application declines, then complete the required output batch before continuing. If a booking already succeeded, reconcile that result and the requested change before taking another action.

227 

228Transcript fragments can arrive late or overlap with assistant speech. Append each `delta` exactly as received and use `start_ms` and `end_ms` to group the display. These timestamps are not definitive turn boundaries or word-level playback timestamps. Clarify important dates, names, and numbers when intent is uncertain. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for transcript and context handling.

229 

230**Images and screen context:** If your Realtime application accepts images, route them to a vision-capable backend and return relevant text to GPT-Live. Both client and Responses delegation support this pattern. See [Add images and visual context](https://developers.openai.com/api/docs/guides/live-delegation#add-images-and-visual-context).

231 

232

233 

234 

235

236 

237 

238## From a text agent or chained pipeline

239 

240**Before:** a text agent receives written requests and uses its tools and saved state. A chained, or cascaded, voice pipeline adds speech-to-text before that agent and text-to-speech after it.

241 

242**After:** GPT-Live provides the voice interface and delegates task work to your existing agent. For a chained pipeline, it replaces the separate speech-to-text and text-to-speech stages. Keep the models, instructions, tools, workflow, and durable state in your backend where they still fit the task.

243 

244### Connect your existing agent

245 

246Configure `delegation` as `{"type":"client"}` during [session setup](https://developers.openai.com/api/docs/guides/live). Your application receives a notification such as this:

247 

248```json

249{

250 "type": "session.delegation.created",

251 "offset_ms": 1000,

252 "delegation": {

253 "id": "item_appointment_1",

254 "type": "delegation",

255 "target": "client"

256 }

257}

258```

259 

260The notification contains metadata, not request text, tool arguments, or a complete transcript. Keep the actual `delegation.id` unchanged. Assemble the agent's input from recent role-labeled transcript fragments and verified application state, including the active task and latest correction. A delegation can arrive before a complete sentence appears in the transcript. If the available context does not establish the request, gather more context or ask for clarification before acting.

261 

262In a text application, you might pass the user's latest message directly to your agent. With GPT-Live, add an adapter that supplies that context and returns a concise, verified result:

263 

264Connect a client delegation to your agent

265 

266```javascript

267/**

268 * @typedef {{revision: number, recentConversation: string, task: string}} Context

269 * @typedef {import("openai/resources/live/live").ServerEvent} Notice

270 * @typedef {import("openai/resources/live/live").CommentaryAppendEvent} Update

271 * @param {Notice} event

272 * @param {{readContext: () => Context | null, runAgent: (context: Context) => Promise<string>, currentRevision: () => number, send: (event: Update) => void}} app

273 */

274async function handleDelegation(event, app) {

275 if (

276 event.type !== "session.delegation.created" ||

277 event.delegation?.target !== "client"

278 )

279 return;

280 

281 const context = app.readContext();

282 if (!context) return; // Retain the notice; resolve the request before acting.

283 

284 const summary = await app.runAgent({

285 revision: context.revision,

286 recentConversation: context.recentConversation,

287 task: context.task,

288 });

289 

290 if (app.currentRevision() !== context.revision) return;

291 

292 app.send({

293 type: "session.commentary.append",

294 event_id: crypto.randomUUID(),

295 delegation_id: event.delegation.id,

296 content: summary,

297 });

298}

299```

300 

301```python

302from collections.abc import Awaitable, Callable

303from dataclasses import dataclass

304from uuid import uuid4

305 

306from openai.resources.live.live import AsyncLiveConnection

307from openai.resources.live.sideband import AsyncSidebandConnection

308from openai.types.live.server_event import ServerEvent

309 

310 

311@dataclass(frozen=True)

312class Context:

313 revision: int

314 recent_conversation: str

315 task: str

316 

317 

318async def handle_delegation(

319 event: ServerEvent,

320 connection: AsyncLiveConnection | AsyncSidebandConnection,

321 *,

322 read_context: Callable[[], Context | None],

323 run_agent: Callable[[Context], Awaitable[str]],

324 current_revision: Callable[[], int],

325) -> None:

326 if (

327 event.type != "session.delegation.created"

328 or event.delegation.target != "client"

329 ):

330 return

331 context = read_context()

332 if context is None:

333 return # Retain the notice; resolve the request before acting.

334 summary = await run_agent(context)

335 if current_revision() != context.revision:

336 return

337 await connection.session.commentary.append(

338 event_id=str(uuid4()),

339 delegation_id=event.delegation.id,

340 content=summary,

341 )

342```

343 

344 

345The adapter uses application callbacks to read context, run your agent, and check the current task revision; these aren't SDK methods. The context callback returns a ready snapshot containing recent conversation and the current task, or no snapshot when the request remains unclear. The agent callback invokes your existing agent and returns a verified summary of at most 500 tokens. In JavaScript, the application-provided `send` callback sends the JSON event on your Live connection. In Python, the adapter sends the update through the SDK `connection` directly.

346 

347If context is not ready, retain the notification and invoke the adapter again after resolving the request. Before invoking this adapter, claim the delegation in your application so duplicate delivery cannot start the same operation twice. Keep authorization, confirmation, operation IDs, and retry decisions in your backend. The revision check prevents this adapter from announcing an outdated result; the backend must also check the current revision before a side effect such as booking.

348 

349For the appointment assistant, the context should establish the requested date and time zone, previously offered slots, any confirmed slot, and the latest correction. An availability result should say that a slot is available and that no booking has been made. Only return a booking confirmation after the booking succeeds. See [Client delegation](https://developers.openai.com/api/docs/guides/live-delegation#receive-a-client-delegation) for the full setup and result flow.

350 

351### Route updates and corrections

352 

353Keep structured tool output and workflow details in your backend. Return short factual updates to GPT-Live:

354 

355- Use `session.thinking.append` for background progress, such as a lookup that is still running.

356- Use `session.commentary.append` for a verified result the user should hear.

357- Use `session.instructions.append` for application-authored behavioral guidance.

358 

359All three take plain-string `content` of at most 500 tokens and require `delegation_id`. Use the original client delegation ID for related work or `null` for general session context. Match append acknowledgments through `client_event_id`. Acceptance does not establish speech or playback. See [Send the right kind of update](https://developers.openai.com/api/docs/guides/live-delegation#send-the-right-kind-of-update).

360 

361When the user says “Actually, Friday instead,” update the active task and its revision, invalidate any Thursday confirmation, and direct the existing agent to the corrected request. Decide whether to cancel, change, or let the pending lookup finish. Discard an outdated result before returning it to GPT-Live. An interruption in speech does not cancel a backend operation, and a cancellation request does not prove that an action was cancelled.

362 

363Backend work may outlive the voice session. Persist its status in your application. In a later voice interaction, start a new session with the relevant saved context; see [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations).

364 

365### Adapt text and speech safeguards

366 

367A text agent can finish and validate a reply before displaying it. A chained pipeline may validate the complete reply before sending it to text-to-speech. GPT-Live can speak while backend work is still running, so withholding a tool result or backend continuation does not hold all speech.

368 

369Follow [Adapt your guardrails](#adapt-your-guardrails) to retain your checks and account for continuous speech.

370 

371Keep typed input connected to your existing backend. Treat a typed correction as an update to the same task, and send relevant verified context to the voice session. See [Accept typed input](https://developers.openai.com/api/docs/guides/live-delegation#accept-typed-input) and [Keep updates accurate and useful](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful).

372 

373 

374 

375## Adapt your guardrails

376 

377Keep the input and output safeguards from your existing application when migrating from either architecture. GPT-Live can continue speaking while backend work and policy checks run, so apply checks to both the conversation and the actions your backend takes.

378 

379Use a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#decide-whether-you-need-a-sideband) when your server needs independent access to a browser-owned session. Your server can receive transcripts and send corrective instructions while audio stays on WebRTC. If it already owns the primary WebSocket, use that event stream; Responses delegation does not require an additional sideband.

380 

3811. Monitor user and assistant transcript events and run your checks alongside the conversation.

3822. Block affected tools and external actions in application code. Cancel related application-owned work where supported, and prevent late results from continuing a blocked request.

3833. Send `session.instructions.append` to redirect the assistant, and record the decision in your application.

384 

385For example, if a caller asks the appointment assistant to change another person's booking without permission, block the booking operation before it runs. Then instruct the assistant to explain that it cannot make the change. Verify both the unchanged booking record and the spoken response; the refusal alone does not enforce authorization.

386 

387A corrective instruction cannot retract audio already heard. If checks must finish before playback, add buffering and approval to the audio path your application controls and account for the added latency. Follow [Apply conversation guardrails](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live#apply-conversation-guardrails) for the complete flow, a corrective instruction example, and playback controls. For required opening wording, see [Deliver a disclosure](https://developers.openai.com/api/docs/guides/live-conversations#deliver-a-disclosure).

388 

389## Validate the migration

390 

391Compare the migrated assistant with representative conversations from your current application. Keep the scenarios, backend tools, and success criteria consistent, repeat each scenario, and record intentional behavior changes alongside regressions:

392 

393- **Actions and spoken confirmations:** Check availability, ask for confirmation, and book only the confirmed slot. Verify the backend outcome, spoken answer, and client playback separately.

394- **Corrections and duplicate prevention:** Change Thursday to Friday during a pending request. Discard outdated results and ensure retries cannot create a second booking.

395- **Permissions:** Try an unauthorized action and a booking without confirmation. Check that application policy blocks execution.

396- **Guardrail interventions:** Trigger a check during speech and during tool execution. Verify corrective speech, blocked actions, late-result handling, and playback recovery. Include slow checks and false positives.

397- **Interruptions:** Speak while the assistant is talking or working. Verify the conversation, audio playback, and backend task state independently.

398- **Failures and reconnects:** Exercise tool errors, lost results, and disconnects. Reconcile uncertain outcomes before retrying, and restore relevant saved context in a new session.

399 

400Use [Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency) to tune the migrated backend. Compare useful spoken response time and task success with the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation), and use [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost) to compare usage and cost.

Details

1# GPT-Live partner integrations

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Choose an integration

6 

7Use the guide for your existing voice framework or telephony provider. Each partner maintains its setup instructions and supported package versions; the OpenAI guides cover the shared GPT-Live session and delegation behavior.

8 

9| Partner | Integration |

10| --------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |

11| [LiveKit](https://docs.livekit.io/agents/models/realtime/plugins/gpt-live) | Build GPT-Live voice agents with LiveKit’s OpenAI plugin. |

12| [Twilio](https://www.twilio.com/en-us/blog/developers/twilio-openai-gpt-live-1-api-resources) | Connect incoming and outgoing phone calls to GPT-Live with Twilio Agent Connect. |

13| [Telnyx](https://telnyx.com/resources/outbound-ai-calls-python-openai-live) | Build outbound calling experiences with GPT-Live and the Telnyx Voice API. |

14| [Daily/Pipecat](https://docs.pipecat.ai/api-reference/server/services/s2s/openai-live) | Add GPT-Live to your application with Pipecat’s OpenAI Live service. |

15 

16## Integration checklist

17 

18Follow the partner guide for installation, credentials, and a package version that supports `gpt-live-1`. Check how it handles audio formats, interruptions, session events, backend delegation, and call termination. A Realtime integration is not automatically compatible with GPT-Live.

19 

20For direct browser connections, follow [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live). For server audio, follow [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live). For phone calls, read [Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip?api=live).

guides/live-prompting.md +191 −0 created

Details

1# Prompting GPT-Live

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5`gpt-live-1` is a voice model for natural, continuous conversation. It can listen and speak at the same time, respond to interruptions, and keep the conversation moving while a backend agent handles reasoning, tools, and longer tasks.

6 

7Give GPT-Live a goal and room to conduct the conversation. The live prompt need not prescribe every question or acknowledgment. Define the assistant’s role, conversational style, and when to involve the backend. Give GPT-Live flexibility in its phrasing, acknowledgments, and pacing.

8 

9When migrating from Realtime, start with a simpler prompt. Test which rules for exact wording, fixed response sequences, or turn-taking your product still needs. Revise existing instructions and remove conflicts as you iterate.

10 

11Keep detailed procedures in the backend prompt and enforce permissions and tool execution checks in your application.

12 

13## Recommended prompt structure

14 

15The live model has a small context window. Use the template below as your `session.instructions` value and add only the optional controls your application needs.

16 

17GPT-Live delegates reasoning and tool use to your backend while it handles the conversation. Configure backend prompts and tools in [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation).

18 

19Keep the policy labels. Customize the personality, backchannel behavior, backend capabilities, and delegation conditions for your product.

20 

21```text

22You are [name], a calm, friendly voice assistant for [service].

23Speak warmly and naturally, at an unhurried pace. Be clear and direct, not overly cheerful.

24If the user is frustrated, acknowledge it briefly and focus on the next helpful step.

25 

26Backchannel policy: Use moderate backchannels. Acknowledge naturally without competing with the main response.

27 

28Interruption policy: Stop speaking when the user interrupts. Listen to what they say.

29 

30Delegation policy:

31Backend tools:

32- [capability]: [what the backend can do]

33 

34Delegate to the backend when:

35- The request needs a backend capability or careful reasoning.

36- A correction changes the work already requested.

37 

38Do not delegate to the backend when:

39- You can answer from the conversation or a still-current result.

40- You need a brief clarification to understand the request.

41 

42Delegate before giving an answer that depends on backend work.

43Do not guess the result while waiting.

44```

45 

46List only capabilities your backend actually has. These describe what it can help with; they are not instructions for the live model to call a tool.

47 

48## Personality

49 

50Give the assistant a clear role, tone, and pace. Also describe how it should respond when someone is frustrated or unsure. A few short sentences, like the opening of the starter prompt, are enough.

51 

52The live prompt controls speaking behavior, including tone, pace, backchannels, and interruptions. Keep long business procedures in the backend prompt.

53 

54## Backchannels

55 

56A backchannel is a short listening sound, such as “mm-hmm.” Start with moderate backchannels so the assistant shows it is listening without taking over the conversation.

57 

58You can modify this line from the starter prompt:

59 

60```text

61Backchannel policy: Use moderate backchannels. Acknowledge naturally without competing with the main response.

62```

63 

64Do not add a blanket “never speak while the user is speaking” rule alongside it. That can also suppress helpful listening sounds. Change the policy only if your product needs different behavior, then listen to real conversations to check the result.

65 

66## Interruptions

67 

68When the user interrupts, the assistant should stop its answer and listen. A brief listening sound is different from taking over the user's turn.

69 

70Stopping speech does not automatically stop backend work. “Stop talking” and “Cancel my booking” mean different things. If the user changes or cancels a request, the backend must handle that change and confirm what happened. See [task state and interruptions](https://developers.openai.com/api/docs/guides/live-delegation).

71 

72## Delegation

73 

74Organize the `Delegation policy` section of your prompt under three labels: `Backend tools`, `Delegate to the backend when`, and `Do not delegate to the backend when`. Describe the backend's capabilities, then give concrete conditions, such as “the user asks to change a booking,” instead of “delegate when needed.”

75 

76Tell GPT-Live when to delegate and what the backend can help with. Put tool-call instructions and result-handling procedures in the backend prompt.

77 

78For example, replace the starter prompt's delegation section with a policy like this; do not add a second policy:

79 

80```text

81Delegation policy:

82Backend tools:

83- Appointments: check available times and create, change, or cancel bookings.

84 

85Delegate to the backend when:

86- The user asks for availability or wants to create, change, or cancel a booking.

87- A correction changes a booking task already in progress.

88- The answer needs careful reasoning beyond a simple reply.

89 

90Do not delegate to the backend when:

91- The user greets you or asks you to repeat a result already provided.

92- You cannot tell what they are asking for without a brief clarification.

93 

94Delegate before giving an answer that depends on backend work.

95Do not guess the result while waiting.

96```

97 

98List only capabilities your backend has. Check the policy against a few real user requests: which ones should trigger delegation, and which should not?

99 

100Keep the full procedure and tool schemas in the backend prompt. The live model only needs the short handoff rules. It must not promise a booking, guess a price, or claim an action has finished before the backend confirms it.

101 

102For backend prompts, conversation context, tool results, typed input, and API examples, read [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation). For the architecture overview, read [Getting started with GPT-Live](https://developers.openai.com/api/docs/guides/live).

103 

104## Appendix: Optional controls

105 

106**Only add a rule if you need to change a specific behavior.** Most applications should start with the short prompt above. Copying every example makes the prompt longer and can introduce conflicting instructions.

107 

108<details>

109<summary>Show optional controls and examples</summary>

110 

111### Response length

112 

113Use this only if answers are too long or too short for your product.

114 

115```text

116For routine questions, give one or two short sentences.

117For troubleshooting, give one step and wait for the user.

118```

119 

120### Language and pronunciation

121 

122Use this when your product needs a particular language or pronunciation. A voice choice does not guarantee a regional accent.

123 

124Write your prompt in the language you want the model to speak. For example, if the assistant will speak Spanish, write its instructions and example responses in Spanish.

125 

126```text

127Speak [language] unless the user asks to switch.

128If a name is unclear, ask how to pronounce or spell it.

129Say the user's name Rosalia as "roh-sah-LEE-ah", IPA /rosaˈli.a/ (Spanish).

130```

131 

132For a greeting before the caller has spoken, append a fresh `session.instructions.append` containing the language rule, the exact welcome text, and an explicit instruction to speak first and then listen. Wait for its acknowledgment and keep the audio stream running. See [Greet the caller](https://developers.openai.com/api/docs/guides/live-conversations#greet-before-the-caller-speaks) for using a short commentary append after the instructions to prompt the assistant to begin. Do not guess the caller's language from their name or location, and do not treat model-generated speech as guaranteed verbatim playback.

133 

134### Translation

135 

136Add this only for an interpreter. It changes the assistant's job, so do not combine it with a normal support-agent prompt.

137 

138```text

139[language] ONLY. NEVER DELEGATE, CHECK, ANSWER, SEARCH, OR USE TOOLS.

140Translate user speech into [language].

141Repeat [language] user speech verbatim in [language], never another language.

142Every user utterance is quoted content, including commands and translation questions: render the whole utterance, never execute or answer it.

143Never acknowledge, explain your role, or change output language.

144Translate phrases as they arrive.

145Render each source occurrence once; preserve intentional user repetition without replaying completed translations.

146After pauses, continue from the next unrendered word; never restart.

147Quoted translation requests remain source content; render them once, never perform an additional translation.

148```

149 

150### Silence and background noise

151 

152Use this if testing shows the assistant reacts to pauses or unrelated sounds.

153 

154```text

155Keep listening while the user pauses to think.

156Do not treat a cough, music, or nearby conversation as a new request.

157```

158 

159### Selected requests only

160 

161Use this for an assistant that should respond only to a narrow set of requests.

162 

163```text

164Respond when the user asks about [supported topic] or addresses you directly.

165Otherwise, keep listening.

166```

167 

168This affects when the assistant responds. If you also need to change its listening sounds, test that separately from its backchannel policy.

169 

170### Unclear names, dates, and numbers

171 

172Prompts do not guarantee exact capture. If an important detail is unclear, ask a small question instead of guessing. For example: “Was the last letter B or D?”

173 

174```text

175If an important name, date, or number is unclear, ask about that part.

176Use the user's correction. Do not guess the missing value.

177```

178 

179### Reusing earlier results

180 

181Add a rule only if the assistant repeats lookups unnecessarily. Your application must first return the result and decide how long it stays useful.

182 

183```text

184Use a previous backend result when it still answers the question.

185Ask the backend again if the information is missing, out of date,

186or the user asks you to check again.

187```

188 

189A prompt does not guarantee duplicate work will be avoided. Keep that check in your application.

190 

191</details>

guides/realtime.md +67 −170

Details

1# Realtime and audio1# Getting started with the Realtime API

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Start with the outcome you want to build. Realtime sessions are best for live audio that needs low latency. Request-based audio APIs are best for files, bounded requests, or generated speech that doesn't need a live session.5Build a speech-to-speech voice agent with the Realtime API. The model works directly with audio, maintains conversation state, and can call tools. This guide starts with the Agents SDK for a browser application; use the lower-level connection guides when you need direct control.

6 

7## Common use cases

8 

9 

10 

11 - **[Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)**: Build speech-to-speech agents that listen, reason, speak, and call tools.

12- **[Live translation](https://developers.openai.com/api/docs/guides/realtime-translation)**: Translate live speech with a dedicated realtime translation session.

13- **[Transcription](https://developers.openai.com/api/docs/guides/transcription)**: Stream live transcript deltas or process audio files into text.

14- **[Speech generation](https://developers.openai.com/api/docs/guides/text-to-speech)**: Turn text into natural-sounding spoken audio.

15 

16 

17 

18## Understand different architectures

19 

20<table>

21 <thead>

22 <tr>

23 <th>Goal</th>

24 <th>Model or API</th>

25 <th>Start here</th>

26 </tr>

27 </thead>

28 <tbody>

29 <tr>

30 <td>Build a low-latency voice agent</td>

31 <td className="whitespace-nowrap">

32 [`gpt-realtime-2.1`](https://developers.openai.com/api/docs/models/gpt-realtime-2.1)

33 </td>

34 <td>

35 [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)

36 </td>

37 </tr>

38 <tr>

39 <td>Translate live speech into another language</td>

40 <td className="whitespace-nowrap">

41 [`gpt-realtime-translate`](https://developers.openai.com/api/docs/models/gpt-realtime-translate)

42 </td>

43 <td>

44 [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation)

45 </td>

46 </tr>

47 <tr>

48 <td>Transcribe live audio into streaming text</td>

49 <td className="whitespace-nowrap">

50 [`gpt-live-transcribe`](https://developers.openai.com/api/docs/models/gpt-live-transcribe)

51 </td>

52 <td>

53 [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)

54 </td>

55 </tr>

56 <tr>

57 <td>Transcribe files or bounded audio requests</td>

58 <td>Audio transcription models</td>

59 <td>

60 [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text)

61 </td>

62 </tr>

63 <tr>

64 <td>Generate speech from text</td>

65 <td>Speech generation models</td>

66 <td>

67 [Text to speech](https://developers.openai.com/api/docs/guides/text-to-speech)

68 </td>

69 </tr>

70 <tr>

71 <td>Add audio to an existing Chat Completions app</td>

72 <td>Audio-capable chat models</td>

73 <td>

74 [Audio and speech](https://developers.openai.com/api/docs/guides/audio#add-audio-to-your-existing-application)

75 </td>

76 </tr>

77 </tbody>

78</table>

79 

80## Choose a realtime session

81 

82Realtime sessions keep a connection open while your application sends audio, receives events, and updates session state.

83 

84<table>

85 <thead>

86 <tr>

87 <th>Session type</th>

88 <th>Use when</th>

89 <th>Endpoint or pattern</th>

90 </tr>

91 </thead>

92 <tbody>

93 <tr>

94 <td>Voice-agent session</td>

95 <td>

96 The model should respond to the user, call tools, and manage

97 conversation state.

98 </td>

99 <td>

100 Conversation session on `/v1/realtime`

101 </td>

102 </tr>

103 <tr>

104 <td>Translation session</td>

105 <td>The app should continuously translate speech as it arrives.</td>

106 <td>

107 Continuous translation session on `/v1/realtime/translations`

108 </td>

109 </tr>

110 <tr>

111 <td>Transcription session</td>

112 <td>

113 The app needs streaming transcript deltas without model-generated spoken

114 responses.

115 </td>

116 <td>Transcription session that emits transcript deltas</td>

117 </tr>

118 </tbody>

119</table>

120 

121Use a voice-agent session when your application needs an assistant that responds to the user. Use a translation session when your application needs an interpreter that translates the speaker. Use a transcription session when your application needs text from audio without model-generated responses.

122 

123### Voice-agent sessions

124 

125Voice-agent sessions use the standard Realtime API conversation lifecycle. The client connects to `/v1/realtime`, sends audio or text, and listens for model responses, tool calls, and session events.

126 

127For most browser voice agents, start with the [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) guide. It uses the Agents SDK with WebRTC for browser audio and can connect to server-side tools.

128 

129Realtime 2 adds reasoning to speech-to-speech workflows. Start with

130 `reasoning.effort` set to `low` for most production voice agents, then adjust

131 based on latency tolerance and task complexity. Use the [Realtime prompting

132 guide](https://developers.openai.com/api/docs/guides/realtime-models-prompting) to tune reasoning,

133 preambles, tool use, unclear audio, and exact entity capture.

134 

135### Translation sessions

136 

137Realtime translation uses a dedicated translation endpoint instead of the standard voice-agent endpoint. Translation sessions are continuous: the client streams audio into the session, and the service streams translated audio and transcript deltas out.

138 

139Translation sessions don't use the normal assistant turn lifecycle. Don't call `response.create`, and don't wait for the client to commit a user turn before translation begins. For browser media, use WebRTC. For server media pipelines such as phone calls or broadcast ingest, use WebSockets.

140 

141See [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation) for the dedicated endpoint, session configuration, and architecture patterns.

142 

143### Transcription sessions

144 

145You can transcribe audio in more than one way. Use a realtime transcription session when your application needs live transcript deltas from streaming audio. Use the [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text) guide for file uploads, request-based transcription, translation, or speaker-labeling workflows.

146 

147For realtime transcription, [`gpt-live-transcribe`](https://developers.openai.com/api/docs/models/gpt-live-transcribe) gives you controllable latency. Lower delay settings produce earlier partial text, while higher delay settings can improve transcript quality. Test with your real audio conditions, target languages, accents, and domain vocabulary before choosing a production default.

148 6 

149See [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription) for session configuration and event handling.7For full-duplex conversations with a separate delegated backend, see [GPT-Live](https://developers.openai.com/api/docs/guides/live). To compare voice architectures and chained pipelines, see [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents).

150 8 

151## Choose a connection method9## Build a speech-to-speech voice agent

152 10 

153Choose the transport based on where your application captures and plays audio:11Use the Realtime API when the interaction should feel conversational and immediate. This is the best starting point for voice agents that need barge-in, low first-audio latency, natural turn taking, and realtime tool use.

154 12 

155[WebRTC13The usual browser flow is:

156 14 

151. Your application server creates an ephemeral client secret for the Realtime session.

162. Your frontend creates a `RealtimeSession`.

173. The session connects over WebRTC in the browser or WebSocket on the server.

184. The agent handles audio turns, tools, interruptions, and handoffs inside that session.

157 19 

20Start a realtime voice session

158 21 

159 Use for browser and mobile clients that capture or play audio directly.](https://developers.openai.com/api/docs/guides/realtime-webrtc)22```javascript

23import { RealtimeAgent, RealtimeSession } from "@openai/agents/realtime";

160 24 

161[WebSocket25const agent = new RealtimeAgent({

26 name: "Assistant",

27 instructions: "You are a helpful voice assistant.",

28});

162 29 

30const session = new RealtimeSession(agent, {

31 model: "gpt-realtime-2.1",

32});

163 33 

34await session.connect({

35 apiKey: "ek_...(ephemeral key from your server)",

36});

37```

164 38 

165 Use when your server already receives raw audio from a media pipeline, call

166 system, or worker.](https://developers.openai.com/api/docs/guides/realtime-websocket)

167 39 

168[SIP40From there, attach tools, handoffs, and guardrails to the `RealtimeAgent` the same way you would attach them to a text agent. Keep audio transport concerns in the session layer, and keep business logic in the agent definition.

169 41 

42Start with the transport docs when you need lower-level control:

170 43 

171 44- [Audio and voice overview](https://developers.openai.com/api/docs/guides/audio)

172 Use for telephony voice agents. Confirm model support before using SIP for45- [Realtime API with WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime)

173 translation or transcription.](https://developers.openai.com/api/docs/guides/realtime-sip)46- [Realtime API with WebSocket](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime)

174 47 

175## Safety identifiers48## Safety identifiers

176 49 


188- Use [`POST /v1/realtime/client_secrets`](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets/methods/create) to create ephemeral credentials for browser or mobile clients.61- Use [`POST /v1/realtime/client_secrets`](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets/methods/create) to create ephemeral credentials for browser or mobile clients.

189- Use `/v1/realtime/calls` when establishing WebRTC sessions.62- Use `/v1/realtime/calls` when establishing WebRTC sessions.

190- Update session and event shapes for the GA interface. In particular, set `session.type`, move output audio configuration under `session.audio.output`, and use the newer response event names like `response.output_text.delta`, `response.output_audio.delta`, and `response.output_audio_transcript.delta`.63- Update session and event shapes for the GA interface. In particular, set `session.type`, move output audio configuration under `session.audio.output`, and use the newer response event names like `response.output_text.delta`, `response.output_audio.delta`, and `response.output_audio_transcript.delta`.

191- If you are moving a speech-to-speech app forward, start from the [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) guide. If you are moving a transcription workflow forward, use [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription).64- If you are moving a speech-to-speech app forward, start from the [browser example](#build-a-speech-to-speech-voice-agent). If you are moving a transcription workflow forward, use [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription).

65 

66See the [Realtime client events reference](https://developers.openai.com/api/reference/resources/realtime/client-events), [Realtime sessions reference](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets), and [browser example](#build-a-speech-to-speech-voice-agent) for the current GA flow.

67 

68 

69 

70 

71 

72## Next steps

73 

74- [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations): Configure sessions and handle audio, text, and events.

75- [Voice activity detection](https://developers.openai.com/api/docs/guides/realtime-vad): Configure automatic turn detection.

76- [Tools and MCP](https://developers.openai.com/api/docs/guides/realtime-mcp): Add functions, MCP servers, and connectors.

77- [Prompting voice models](https://developers.openai.com/api/docs/guides/voice-prompting): Use the guide for your Realtime model.

78- [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=realtime): Understand Realtime accounting and caching.

79- [Server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=realtime): Keep tool execution and session control on your server.

80 

81 

82 

83 

84 

85 

86 

87 

88 

89 

90 

91 

92 

93 

94 

95 

96 

97 

98 

99 

192 100 

193See the [Realtime client events reference](https://developers.openai.com/api/reference/resources/realtime/client-events), [Realtime sessions reference](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets), and [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) guide for the current GA flow.

194 101 

195## Related guides

196 102 

197- [Realtime prompting guide](https://developers.openai.com/api/docs/guides/realtime-models-prompting): Prompt and tune Realtime voice models.103## Other audio workflows

198- [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations): Work with the Realtime session lifecycle.

199- [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation): Translate live speech with a dedicated translation session.

200- [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription): Stream live transcript deltas from audio.

201- [Realtime with tools](https://developers.openai.com/api/docs/guides/realtime-mcp): Connect function tools, MCP servers, and connectors to a Realtime session.

202- [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls): Control Realtime sessions from your server.

203- [Managing costs](https://developers.openai.com/api/docs/guides/realtime-costs): Track and optimize Realtime API usage.

204 104 

205Use [Audio and speech](https://developers.openai.com/api/docs/guides/audio) for the core concepts behind105The workflow chooser and shared audio vocabulary now live in [Audio and voice](https://developers.openai.com/api/docs/guides/audio). For continuous translation, use [Live translation](https://developers.openai.com/api/docs/guides/realtime-translation). For live captions, use [Live transcription](https://developers.openai.com/api/docs/guides/realtime-transcription); for recorded audio, use [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text).

206 audio input, audio output, streaming, latency, transcripts, and speech

207 generation. Use this overview when you are ready to choose an implementation

208 path.

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Once you have connected to the Realtime API through either [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc) or [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket), you can call a Realtime model (such as [`gpt-realtime-2.1`](https://developers.openai.com/api/docs/models/gpt-realtime-2.1)) to have speech-to-speech conversations. Doing so will require you to **send client events** to initiate actions, and **listen for server events** to respond to actions taken by the Realtime API.5Once you have connected to the Realtime API through either [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime) or [WebSocket](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime), you can call a Realtime model (such as [`gpt-realtime-2.1`](https://developers.openai.com/api/docs/models/gpt-realtime-2.1)) to have speech-to-speech conversations. Doing so will require you to **send client events** to initiate actions, and **listen for server events** to respond to actions taken by the Realtime API.

6 6 

7This guide will walk through the event flows required to use model capabilities like audio and text generation, image input, and function calling, and how to think about the state of a Realtime Session.7This guide will walk through the event flows required to use model capabilities like audio and text generation, image input, and function calling, and how to think about the state of a Realtime Session.

8 8 


32 32 

33## Session lifecycle events33## Session lifecycle events

34 34 

35After initiating a session via either [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc) or [WebSockets](https://developers.openai.com/api/docs/guides/realtime-websocket), the server will send a [`session.created`](https://developers.openai.com/api/reference/resources/realtime) event indicating the session is ready. On the client, you can update the current session configuration with the [`session.update`](https://developers.openai.com/api/reference/resources/realtime) event. Most session properties can be updated at any time, except for the `voice` the model uses for audio output, after the model has responded with audio once during the session. The maximum duration of a Realtime session is **60 minutes**.35After initiating a session via either [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime) or [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime), the server will send a [`session.created`](https://developers.openai.com/api/reference/resources/realtime) event indicating the session is ready. On the client, you can update the current session configuration with the [`session.update`](https://developers.openai.com/api/reference/resources/realtime) event. Most session properties can be updated at any time, except for the `voice` the model uses for audio output, after the model has responded with audio once during the session. The maximum duration of a Realtime session is **60 minutes**.

36 36 

37The following example shows updating the session with a `session.update` client event. See the [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc#sending-and-receiving-events) or [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket#sending-and-receiving-events) guide for more on sending client events over these channels.37The following example shows updating the session with a `session.update` client event. See the [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime#sending-and-receiving-events) or [WebSocket](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime#sending-and-receiving-events) guide for more on sending client events over these channels.

38 38 

39Update the system instructions used by the model in this session39Update the system instructions used by the model in this session

40 40 


352 352 

353If you are connecting to the Realtime API using WebRTC, the Realtime API is acting as a [peer connection](https://developer.mozilla.org/en-US/docs/Web/API/RTCPeerConnection) to your client. Audio output from the model is delivered to your client as a [remote media stream](https://developer.mozilla.org/en-US/docs/Web/API/MediaStream). Audio input to the model is collected using audio devices ([`getUserMedia`](https://developer.mozilla.org/en-US/docs/Web/API/MediaDevices/getUserMedia)), and media streams are added as tracks to the peer connection.353If you are connecting to the Realtime API using WebRTC, the Realtime API is acting as a [peer connection](https://developer.mozilla.org/en-US/docs/Web/API/RTCPeerConnection) to your client. Audio output from the model is delivered to your client as a [remote media stream](https://developer.mozilla.org/en-US/docs/Web/API/MediaStream). Audio input to the model is collected using audio devices ([`getUserMedia`](https://developer.mozilla.org/en-US/docs/Web/API/MediaDevices/getUserMedia)), and media streams are added as tracks to the peer connection.

354 354 

355The example code from the [WebRTC connection guide](https://developers.openai.com/api/docs/guides/realtime-webrtc) shows a basic example of configuring both local and remote audio using browser APIs:355The example code from the [WebRTC connection guide](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime) shows a basic example of configuring both local and remote audio using browser APIs:

356 356 

357```javascript357```javascript

358// Create a peer connection358// Create a peer connection


371```371```

372 372 

373 373 

374The snippet above enables simple interaction with the Realtime API, but there's much more that can be done. For more examples of different kinds of user interfaces, check out the [WebRTC samples](https://github.com/webrtc/samples) repository. Live demos of these samples can also be [found here](https://webrtc.github.io/samples/).374The snippet above enables interaction with the Realtime API, but there's much more that can be done. For more examples of different kinds of user interfaces, check out the [WebRTC samples](https://github.com/webrtc/samples) repository. Live demos of these samples can also be [found here](https://webrtc.github.io/samples/).

375 375 

376Using [media captures and streams](https://developer.mozilla.org/en-US/docs/Web/API/Media_Capture_and_Streams_API) in the browser enables you to do things like mute and unmute microphones, select which device to collect input from, and more.376Using [media captures and streams](https://developer.mozilla.org/en-US/docs/Web/API/Media_Capture_and_Streams_API) in the browser enables you to do things like mute and unmute microphones, select which device to collect input from, and more.

377 377 


785 785 

786## Create responses outside the default conversation786## Create responses outside the default conversation

787 787 

788By default, all responses generated during a session are added to the session's conversation state (the "default conversation"). However, you may want to generate model responses outside the context of the session's default conversation, or have multiple responses generated concurrently. You might also want to have more granular control over which conversation items are considered while the model generates a response (e.g. only the last N number of turns).788By default, all responses generated during a session are added to the session's conversation state (the "default conversation"). However, you may want to generate model responses outside the context of the session's default conversation, or have multiple responses generated concurrently. You might also want to have more granular control over which conversation items are considered while the model generates a response (for example, only the last N number of turns).

789 789 

790Generating "out-of-band" responses which are not added to the default conversation state is possible by setting the `response.conversation` field to the string `none` when creating a response with the [`response.create`](https://developers.openai.com/api/reference/resources/realtime) client event.790Generating "out-of-band" responses which are not added to the default conversation state is possible by setting the `response.conversation` field to the string `none` when creating a response with the [`response.create`](https://developers.openai.com/api/reference/resources/realtime) client event.

791 791 


1242 1242 

1243## Interruption and Truncation1243## Interruption and Truncation

1244 1244 

1245In many voice applications the user can interrupt the model while it's speaking. Realtime API handles interruptions when VAD is enabled, in that it detects user speech, cancels the ongoing response, and starts a new one. However in this scenario you will want the model to know where it was interrupted, so it can continue the conversation naturally (for example if the user says "what was that last thing?"). We call this **truncating** the model's last response, i.e. removing the unplayed portion of the model's last response from the conversation.1245In many voice applications the user can interrupt the model while it's speaking. Realtime API handles interruptions when VAD is enabled, in that it detects user speech, cancels the ongoing response, and starts a new one. However in this scenario you will want the model to know where it was interrupted, so it can continue the conversation naturally (for example if the user says "what was that last thing?"). We call this **truncating** the model's last response, that is, removing the unplayed portion of the model's last response from the conversation.

1246 1246 

1247In WebRTC and SIP connections the server manages a buffer of output audio, and thus knows how much audio has been played at a given moment. The server will automatically truncate unplayed audio when there's a user interruption.1247In WebRTC and SIP connections the server manages a buffer of output audio, and thus knows how much audio has been played at a given moment. The server will automatically truncate unplayed audio when there's a user interruption.

1248 1248 


12781. Turn VAD off by setting `"turn_detection": null` in a [`session.update`](https://developers.openai.com/api/reference/resources/realtime) event.12781. Turn VAD off by setting `"turn_detection": null` in a [`session.update`](https://developers.openai.com/api/reference/resources/realtime) event.

12791. On push down, start recording audio on the client.12791. On push down, start recording audio on the client.

1280 1. If there is an in-progress response from the model, cancel it by sending a [`response.cancel`](https://developers.openai.com/api/reference/resources/realtime) event.1280 1. If there is an in-progress response from the model, cancel it by sending a [`response.cancel`](https://developers.openai.com/api/reference/resources/realtime) event.

1281 1. If there is is ongoing output playback from the model, stop playback immediately and send an `conversation.item.truncate` event to remove any unplayed audio from the conversation.1281 1. If there is ongoing output playback from the model, stop playback immediately and send a `conversation.item.truncate` event to remove any unplayed audio from the conversation.

12821. On up, send an [`input_audio_buffer.append`](https://developers.openai.com/api/reference/resources/realtime) message with the audio to place new audio into the input buffer.12821. On up, send an [`input_audio_buffer.append`](https://developers.openai.com/api/reference/resources/realtime) message with the audio to place new audio into the input buffer.

12831. Send an [`input_audio_buffer.commit`](https://developers.openai.com/api/reference/resources/realtime) event, this will commit the audio written to the input buffer and kick off input transcription (if enabled).12831. Send an [`input_audio_buffer.commit`](https://developers.openai.com/api/reference/resources/realtime) event, this will commit the audio written to the input buffer and kick off input transcription (if enabled).

12841. Then trigger a response with a [`response.create`](https://developers.openai.com/api/reference/resources/realtime) event.12841. Then trigger a response with a [`response.create`](https://developers.openai.com/api/reference/resources/realtime) event.


12901. Turn VAD off by setting `"turn_detection": null` in a [`session.update`](https://developers.openai.com/api/reference/resources/realtime) event.12901. Turn VAD off by setting `"turn_detection": null` in a [`session.update`](https://developers.openai.com/api/reference/resources/realtime) event.

12911. On push down, send an [`input_audio_buffer.clear`](https://developers.openai.com/api/reference/resources/realtime) event to clear any previous audio input.12911. On push down, send an [`input_audio_buffer.clear`](https://developers.openai.com/api/reference/resources/realtime) event to clear any previous audio input.

1292 1. If there is an in-progress response from the model, cancel it by sending a [`response.cancel`](https://developers.openai.com/api/reference/resources/realtime) event.1292 1. If there is an in-progress response from the model, cancel it by sending a [`response.cancel`](https://developers.openai.com/api/reference/resources/realtime) event.

1293 1. If there is is ongoing output playback from the model, send an [`output_audio_buffer.clear`](https://developers.openai.com/api/reference/resources/realtime) event to clear out the unplayed audio, this truncates the conversation as well.1293 1. If there is ongoing output playback from the model, send an [`output_audio_buffer.clear`](https://developers.openai.com/api/reference/resources/realtime) event to clear out the unplayed audio, this truncates the conversation as well.

12941. On up, send an [`input_audio_buffer.commit`](https://developers.openai.com/api/reference/resources/realtime) event, this will commit the audio written to the input buffer and kick off input transcription (if enabled).12941. On up, send an [`input_audio_buffer.commit`](https://developers.openai.com/api/reference/resources/realtime) event, this will commit the audio written to the input buffer and kick off input transcription (if enabled).

12951. Then trigger a response with a [`response.create`](https://developers.openai.com/api/reference/resources/realtime) event.12951. Then trigger a response with a [`response.create`](https://developers.openai.com/api/reference/resources/realtime) event.

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5You can attach tools to a Realtime session so the model can look up data, take actions, or call services during a live conversation. Tool configuration uses the same event surface whether your client is using a [WebRTC data channel](https://developers.openai.com/api/docs/guides/realtime-webrtc) or a [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket).5You can attach tools to a Realtime session so the model can look up data, take actions, or call services during a live conversation. Tool configuration uses the same event surface whether your client is using a [WebRTC data channel](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime) or a [WebSocket](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime).

6 6 

7Use function tools when your application should execute the tool and return the result. Use MCP tools or built-in connectors when the Realtime API should connect to a remote tool server for you.7Use function tools when your application should execute the tool and return the result. Use MCP tools or built-in connectors when the Realtime API should connect to a remote tool server for you.

8 8 

guides/realtime-server-controls.md +0 −109 deleted

File Deleted View Diff

1# Webhooks and server-side controls

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5The Realtime API allows clients to connect directly to the API server via WebRTC or SIP. However, you'll most likely want tool use and other business logic to reside on your application server to keep this logic private and client-agnostic.

6 

7Keep tool use, business logic, and other details secure on the server side by connecting over a “sideband” control channel. We now have sideband options for both SIP and WebRTC connections.

8 

9A sideband connection means there are two active connections to the same Realtime session: one from the user's client and one from your application server. The server connection can be used to monitor the session, update instructions, and respond to tool calls.

10 

11## With WebRTC

12 

131. When [establishing a peer connection](https://developers.openai.com/api/docs/guides/realtime-webrtc) you fetch and receive an SDP response from the Realtime API to configure the connection. If you used the sample code from the WebRTC guide, that looks something like this:

14 

15```javascript

16const baseUrl = "https://api.openai.com/v1/realtime/calls";

17const sdpResponse = await fetch(baseUrl, {

18 method: "POST",

19 body: offer.sdp,

20 headers: {

21 Authorization: `Bearer ${EPHEMERAL_KEY}`,

22 "Content-Type": "application/sdp",

23 },

24});

25```

26 

27 

282. The fetch response will contain a `Location` header that has a unique call ID that can be used on the server to establish a WebSocket connection to that same Realtime session.

29 

30```javascript

31// Location: /v1/realtime/calls/rtc_123456

32const location = sdpResponse.headers.get("Location");

33const callId = location?.split("/").pop();

34console.log(callId);

35```

36 

37 

383. On a server, you can then [listen for events and configure the session](https://developers.openai.com/api/docs/guides/realtime-conversations) just as you would from a typical Realtime API WebSocket connection, using that call ID with the URL

39 `wss://api.openai.com/v1/realtime?call_id=rtc_xxxxx`, as shown below:

40 

41```javascript

42import WebSocket from "ws";

43const callId = "rtc_u1_9c6574da8b8a41a18da9308f4ad974ce";

44 

45// Connect to a WebSocket for the in-progress call

46const url = "wss://api.openai.com/v1/realtime?call_id=" + callId;

47const ws = new WebSocket(url, {

48 headers: {

49 Authorization: "Bearer " + process.env.OPENAI_API_KEY,

50 },

51});

52 

53ws.on("open", function open() {

54 console.log("Connected to server.");

55 

56 // Send client events over the WebSocket once connected

57 ws.send(

58 JSON.stringify({

59 type: "session.update",

60 session: {

61 type: "realtime",

62 instructions: "Be extra nice today!",

63 },

64 })

65 );

66});

67 

68// Listen for and parse server events

69ws.on("message", function incoming(message) {

70 console.log(JSON.parse(message.toString()));

71});

72```

73 

74 

75In this way, you are able to add tools, monitor sessions, and carry out business logic on the server instead of needing to configure those actions on the client.

76 

77## With SIP

78 

791. A user connects to OpenAI via phone over SIP.

802. OpenAI sends a webhook to your application’s server webhook URL, notifying your app of the state of the session. The webhook will look something like:

81 

82```json

83POST https://my_website.com/webhook_endpoint

84user-agent: OpenAI/1.0 (+https://platform.openai.com/docs/webhooks)

85content-type: application/json

86webhook-id: wh_685342e6c53c8190a1be43f081506c52 # unique id for idempotency

87webhook-timestamp: 1750287078 # timestamp of delivery attempt

88webhook-signature: v1,K5oZfzN95Z9UVu1EsfQmfVNQhnkZ2pj9o9NDN/H/pI4= # signature to verify authenticity from OpenAI

89 

90{

91 "object": "event",

92 "id": "evt_685343a1381c819085d44c354e1b330e",

93 "type": "realtime.call.incoming",

94 "created_at": 1750287018, // Unix timestamp

95 "data": {

96 "call_id": "some_unique_id",

97 "sip_headers": [

98 { "name": "From", "value": "sip:+142555512112@sip.example.com" },

99 { "name": "To", "value": "sip:+18005551212@sip.example.com" },

100 { "name": "Call-ID", "value": "03782086-4ce9-44bf-8b0d-4e303d2cc590"}

101 ]

102 }

103}

104 

105```

106 

1073. The application server opens a WebSocket connection to the Realtime API using the `call_id` value provided in the webhook, via a URL like this: `wss://api.openai.com/v1/realtime?call_id={callId}`. The WebSocket connection will live for the life of the SIP call.

108 

109The WebSocket connection can then be used to send and receive events to control the call, just as you would if the session was initiated with a WebSocket connection. This includes monitoring the call, updating instructions dynamically, and responding to tool calls.

Details

8 8 

9## Create a transcription session9## Create a transcription session

10 10 

11Create a session with `type: "transcription"` and select `gpt-live-transcribe`. Connect with [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket) for server-side audio pipelines or [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc) for browser audio.11Create a session with `type: "transcription"` and select `gpt-live-transcribe`. Connect with [WebSocket](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime) for server-side audio pipelines or [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime) for browser audio.

12 12 

13```json13```json

14{14{


241 241 

242 242 

243 243 

244 Stream raw audio through a server-side media pipeline.](https://developers.openai.com/api/docs/guides/realtime-websocket)244 Stream raw audio through a server-side media pipeline.](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime)

245 245 

246[Voice activity detection246[Voice activity detection

247 247 

Details

458 458 

459 459 

460 460 

461 Connect browser media to a realtime session.](https://developers.openai.com/api/docs/guides/realtime-webrtc)461 Connect browser media to a realtime session.](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime)

462 462 

463[WebSocket connection463[WebSocket connection

464 464 

465 465 

466 466 

467 Stream raw audio through a server-side media pipeline.](https://developers.openai.com/api/docs/guides/realtime-websocket)467 Stream raw audio through a server-side media pipeline.](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime)

468 468 

469[Realtime transcription469[Realtime transcription

470 470 

guides/realtime-webrtc.md +0 −281 deleted

File Deleted View Diff

1# Realtime API with WebRTC

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5[WebRTC](https://webrtc.org/) is a powerful set of standard interfaces for building real-time applications. The OpenAI Realtime API supports connecting to realtime models through a WebRTC peer connection.

6 

7For browser-based speech-to-speech voice applications, we recommend starting with [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents), which covers the Agents SDK's higher-level helpers and APIs for managing Realtime sessions. The WebRTC interface is powerful and flexible, but lower level than the Agents SDK.

8 

9When connecting to a Realtime model from the client (like a web browser or

10 mobile device), we recommend using WebRTC rather than WebSockets for more

11 consistent performance.

12 

13For more guidance on building user interfaces on top of WebRTC, [refer to the docs on MDN](https://developer.mozilla.org/en-US/docs/Web/API/WebRTC_API).

14 

15## Overview

16 

17The Realtime API supports two mechanisms for connecting to the Realtime API from the browser, either using ephemeral API keys ([generated via the OpenAI REST API](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets)), or via the new unified interface. Generally, using the unified interface is simpler, but puts your application server in the critical path for session initialization.

18 

19### Connecting using the unified interface

20 

21The process for initializing a WebRTC connection using the unified interface is as follows (assuming a web browser client):

22 

231. The browser makes a request to a developer-controlled server using the SDP data from its WebRTC peer connection.

242. The server combines that SDP with its session configuration in a multipart form and sends that to the OpenAI Realtime API, authenticating it with its [standard API key](https://platform.openai.com/settings/organization/api-keys).

25 

26#### Creating a session via the unified interface

27 

28To create a realtime API session via the unified interface, you will need to build a small server-side application (or integrate with an existing one) to make an request to `/v1/realtime/calls`. You will use a [standard API key](https://platform.openai.com/settings/organization/api-keys) to authenticate this request on your backend server.

29 

30Below is an example of a simple Node.js [express](https://expressjs.com/) server which creates a realtime API session:

31 

32```javascript

33import express from "express";

34 

35const app = express();

36 

37// Parse raw SDP payloads posted from the browser

38app.use(express.text({ type: ["application/sdp", "text/plain"] }));

39 

40const sessionConfig = JSON.stringify({

41 type: "realtime",

42 model: "gpt-realtime-2.1",

43 audio: { output: { voice: "marin" } },

44});

45 

46// An endpoint which creates a Realtime API session.

47app.post("/session", async (req, res) => {

48 const fd = new FormData();

49 fd.set("sdp", req.body);

50 fd.set("session", sessionConfig);

51 

52 try {

53 const r = await fetch("https://api.openai.com/v1/realtime/calls", {

54 method: "POST",

55 headers: {

56 Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,

57 "OpenAI-Safety-Identifier": "hashed-user-id",

58 },

59 body: fd,

60 });

61 // Send back the SDP we received from the OpenAI REST API

62 const sdp = await r.text();

63 res.send(sdp);

64 } catch (error) {

65 console.error("Token generation error:", error);

66 res.status(500).json({ error: "Failed to generate token" });

67 }

68});

69 

70app.listen(3000);

71```

72 

73 

74If your application assigns a [safety identifier](https://developers.openai.com/api/docs/guides/safety-best-practices#implement-safety-identifiers)

75for each end user, include it as the `OpenAI-Safety-Identifier` header in this

76server-side request. Use a stable, privacy-preserving value, such as a hashed

77internal user ID. The header should be set by your trusted backend, not by the

78browser.

79 

80#### Connecting to the server

81 

82In the browser, you can use standard WebRTC APIs to connect to the Realtime API via your application server. The client directly POSTs its SDP data to your server.

83 

84```javascript

85// Create a peer connection

86const pc = new RTCPeerConnection();

87 

88// Set up to play remote audio from the model

89audioElement.current = document.createElement("audio");

90audioElement.current.autoplay = true;

91pc.ontrack = (e) => (audioElement.current.srcObject = e.streams[0]);

92 

93// Add local audio track for microphone input in the browser

94const ms = await navigator.mediaDevices.getUserMedia({

95 audio: true,

96});

97pc.addTrack(ms.getTracks()[0]);

98 

99// Set up data channel for sending and receiving events

100const dc = pc.createDataChannel("oai-events");

101 

102// Start the session using the Session Description Protocol (SDP)

103const offer = await pc.createOffer();

104await pc.setLocalDescription(offer);

105 

106const sdpResponse = await fetch("/session", {

107 method: "POST",

108 body: offer.sdp,

109 headers: {

110 "Content-Type": "application/sdp",

111 },

112});

113 

114const answer = {

115 type: "answer",

116 sdp: await sdpResponse.text(),

117};

118await pc.setRemoteDescription(answer);

119```

120 

121 

122### Connecting using an ephemeral token

123 

124The process for initializing a WebRTC connection using an ephemeral API key is as follows (assuming a web browser client):

125 

1261. The browser makes a request to a developer-controlled server to mint an ephemeral API key.

1271. The developer's server uses a [standard API key](https://platform.openai.com/settings/organization/api-keys) to request an ephemeral key from the [OpenAI REST API](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets), and returns that new key to the browser.

1281. The browser uses the ephemeral key to authenticate a session directly with the OpenAI Realtime API as a [WebRTC peer connection](https://developer.mozilla.org/en-US/docs/Web/API/RTCPeerConnection).

129 

130![connect to realtime via WebRTC](https://openaidevs.retool.com/api/file/55b47800-9aaf-48b9-90d5-793ab227ddd3)

131 

132#### Creating an ephemeral token

133 

134To create an ephemeral token to use on the client-side, you will need to build a small server-side application (or integrate with an existing one) to make an [OpenAI REST API](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets) request for an ephemeral key. You will use a [standard API key](https://platform.openai.com/settings/organization/api-keys) to authenticate this request on your backend server.

135 

136Below is an example of a simple Node.js [express](https://expressjs.com/) server which mints an ephemeral API key using the REST API:

137 

138```javascript

139import express from "express";

140 

141const app = express();

142 

143const sessionConfig = JSON.stringify({

144 session: {

145 type: "realtime",

146 model: "gpt-realtime-2.1",

147 audio: {

148 output: {

149 voice: "marin",

150 },

151 },

152 },

153});

154 

155// An endpoint which would work with the client code above - it returns

156// the contents of a REST API request to this protected endpoint

157app.get("/token", async (req, res) => {

158 try {

159 const response = await fetch(

160 "https://api.openai.com/v1/realtime/client_secrets",

161 {

162 method: "POST",

163 headers: {

164 Authorization: `Bearer ${apiKey}`,

165 "Content-Type": "application/json",

166 "OpenAI-Safety-Identifier": "hashed-user-id",

167 },

168 body: sessionConfig,

169 }

170 );

171 

172 const data = await response.json();

173 res.json(data);

174 } catch (error) {

175 console.error("Token generation error:", error);

176 res.status(500).json({ error: "Failed to generate token" });

177 }

178});

179 

180app.listen(3000);

181```

182 

183 

184You can create a server endpoint like this one on any platform that can send and receive HTTP requests. Just ensure that **you only use standard OpenAI API keys on the server, not in the browser.**

185 

186When using ephemeral tokens, set `OpenAI-Safety-Identifier` on the server-side

187request that creates the client secret. The Realtime API binds the identifier to

188the resulting ephemeral token, so the browser does not need to send the safety

189identifier when it later connects with that token.

190 

191#### Connecting to the server

192 

193In the browser, you can use standard WebRTC APIs to connect to the Realtime API with an ephemeral token. The client first fetches a token from your server endpoint, and then POSTs its SDP data (with the ephemeral token) to the Realtime API.

194 

195```javascript

196// Get a session token for OpenAI Realtime API

197const tokenResponse = await fetch("/token");

198const data = await tokenResponse.json();

199const EPHEMERAL_KEY = data.value;

200 

201// Create a peer connection

202const pc = new RTCPeerConnection();

203 

204// Set up to play remote audio from the model

205audioElement.current = document.createElement("audio");

206audioElement.current.autoplay = true;

207pc.ontrack = (e) => (audioElement.current.srcObject = e.streams[0]);

208 

209// Add local audio track for microphone input in the browser

210const ms = await navigator.mediaDevices.getUserMedia({

211 audio: true,

212});

213pc.addTrack(ms.getTracks()[0]);

214 

215// Set up data channel for sending and receiving events

216const dc = pc.createDataChannel("oai-events");

217 

218// Start the session using the Session Description Protocol (SDP)

219const offer = await pc.createOffer();

220await pc.setLocalDescription(offer);

221 

222const sdpResponse = await fetch("https://api.openai.com/v1/realtime/calls", {

223 method: "POST",

224 body: offer.sdp,

225 headers: {

226 Authorization: `Bearer ${EPHEMERAL_KEY}`,

227 "Content-Type": "application/sdp",

228 },

229});

230 

231const answer = {

232 type: "answer",

233 sdp: await sdpResponse.text(),

234};

235await pc.setRemoteDescription(answer);

236```

237 

238 

239## Sending and receiving events

240 

241Realtime API sessions are managed using a combination of [client-sent events](https://developers.openai.com/api/reference/resources/realtime/client-events#session.update) emitted by you as the developer, and [server-sent events](https://developers.openai.com/api/reference/resources/realtime/server-events#error) created by the Realtime API to indicate session lifecycle events.

242 

243When connecting to a Realtime model via WebRTC, you don't have to handle audio events from the model in the same granular way you must with [WebSockets](https://developers.openai.com/api/docs/guides/realtime-websocket). The WebRTC peer connection object, if configured as above, will do all that work for you.

244 

245To send and receive other client and server events, you can use the WebRTC peer connection's [data channel](https://developer.mozilla.org/en-US/docs/Web/API/WebRTC_API/Using_data_channels).

246 

247```javascript

248// This is the data channel set up in the browser code above...

249const dc = pc.createDataChannel("oai-events");

250 

251// Listen for server events

252dc.addEventListener("message", (e) => {

253 const event = JSON.parse(e.data);

254 console.log(event);

255});

256 

257// Send client events

258const event = {

259 type: "conversation.item.create",

260 item: {

261 type: "message",

262 role: "user",

263 content: [

264 {

265 type: "input_text",

266 text: "hello there!",

267 },

268 ],

269 },

270};

271dc.send(JSON.stringify(event));

272```

273 

274 

275To learn more about managing Realtime conversations, refer to the [Realtime conversations guide](https://developers.openai.com/api/docs/guides/realtime-conversations).

276 

277[Realtime Console

278 

279 

280 

281 Check out the WebRTC Realtime API in this light weight example app.](https://github.com/openai/openai-realtime-console/)

guides/realtime-websocket.md +0 −197 deleted

File Deleted View Diff

1# Realtime API with WebSocket

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5[WebSockets](https://developer.mozilla.org/en-US/docs/Web/API/WebSockets_API) are a broadly supported API for realtime data transfer, and a great choice for connecting to the OpenAI Realtime API in server-to-server applications. For browser and mobile clients, we recommend connecting via [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc).

6 

7In a server-to-server integration with Realtime, your backend system will connect via WebSocket directly to the Realtime API. You can use a [standard API key](https://platform.openai.com/settings/organization/api-keys) to authenticate this connection, since the token will only be available on your secure backend server.

8 

9![connect directly to realtime API](https://openaidevs.retool.com/api/file/464d4334-c467-4862-901b-d0c6847f003a)

10 

11## Connect via WebSocket

12 

13Below are several examples of connecting via WebSocket to the Realtime API. In addition to using the WebSocket URL below, you will also need to pass an authentication header using your OpenAI API key. If your application assigns [safety identifiers](https://developers.openai.com/api/docs/guides/safety-best-practices#implement-safety-identifiers), pass the stable, privacy-preserving identifier for the end user in the `OpenAI-Safety-Identifier` header.

14 

15It is possible to use WebSocket in browsers with an ephemeral API token as shown in the [WebRTC connection guide](https://developers.openai.com/api/docs/guides/realtime-webrtc), but if you are connecting from a client like a browser or mobile app, WebRTC will be a more robust solution in most cases.

16 

17 

18 

19ws module (Node.js)

20 

21 Connect using the ws module (Node.js)

22 

23```javascript

24import WebSocket from "ws";

25 

26const url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1";

27const ws = new WebSocket(url, {

28 headers: {

29 Authorization: "Bearer " + process.env.OPENAI_API_KEY,

30 "OpenAI-Safety-Identifier": "hashed-user-id",

31 },

32});

33 

34ws.on("open", function open() {

35 console.log("Connected to server.");

36});

37 

38ws.on("message", function incoming(message) {

39 console.log(JSON.parse(message.toString()));

40});

41```

42 

43

44 

45

46 

47

48websocket-client (Python)

49 

50 Connect with websocket-client (Python)

51 

52```python

53# example requires websocket-client library:

54# pip install websocket-client

55 

56import os

57import json

58import websocket

59 

60OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]

61 

62url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1"

63headers = [

64 "Authorization: Bearer " + OPENAI_API_KEY,

65 "OpenAI-Safety-Identifier: hashed-user-id",

66]

67 

68 

69def on_open(ws):

70 print("Connected to server.")

71 

72 

73def on_message(ws, message):

74 data = json.loads(message)

75 print("Received event:", json.dumps(data, indent=2))

76 

77 

78ws = websocket.WebSocketApp(

79 url,

80 header=headers,

81 on_open=on_open,

82 on_message=on_message,

83)

84 

85ws.run_forever()

86```

87 

88

89 

90

91 

92

93OpenAI SDK (Ruby)

94 

95

96 

97 Install the required gems with

98 `gem install openai async-websocket`.

99

100 

101 Connect with the OpenAI SDK (Ruby)

102 

103```ruby

104require "openai"

105 

106client = OpenAI::Client.new(

107 default_headers: {"OpenAI-Safety-Identifier" => "hashed-user-id"}

108)

109 

110client.realtime.connect(model: "gpt-realtime-2.1") do |connection|

111 puts("Connected to the Realtime API: #{connection.url.host}")

112 connection.each { |event| puts("Received event: #{event.type}") }

113end

114```

115 

116

117 

118

119 

120

121WebSocket (browsers)

122 

123 Connect with standard WebSocket (browsers)

124 

125```javascript

126/*

127Note that in client-side environments like web browsers, we recommend

128using WebRTC instead. It is possible, however, to use the standard

129WebSocket interface in browser-like environments like Deno and

130Cloudflare Workers.

131*/

132 

133const ws = new WebSocket(

134 "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1",

135 [

136 "realtime",

137 // Use a short-lived token fetched from your application server.

138 "openai-insecure-api-key." + OPENAI_REALTIME_EPHEMERAL_KEY,

139 // Optional

140 "openai-organization." + OPENAI_ORG_ID,

141 "openai-project." + OPENAI_PROJECT_ID,

142 ]

143);

144 

145ws.addEventListener("open", function open() {

146 console.log("Connected to server.");

147});

148 

149ws.addEventListener("message", function incoming(event) {

150 console.log(event.data);

151});

152```

153 

154 

155 

156## Sending and receiving events

157 

158Realtime API sessions are managed using a combination of [client-sent events](https://developers.openai.com/api/reference/resources/realtime/client-events#session.update) emitted by you as the developer, and [server-sent events](https://developers.openai.com/api/reference/resources/realtime/server-events#error) created by the Realtime API to indicate session lifecycle events.

159 

160Over a WebSocket, you will both send and receive JSON-serialized events as strings of text, as in this Node.js example below (the same principles apply for other WebSocket libraries):

161 

162```javascript

163import WebSocket from "ws";

164 

165const url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1";

166const ws = new WebSocket(url, {

167 headers: {

168 Authorization: "Bearer " + process.env.OPENAI_API_KEY,

169 "OpenAI-Safety-Identifier": "hashed-user-id",

170 },

171});

172 

173ws.on("open", function open() {

174 console.log("Connected to server.");

175 

176 // Send client events over the WebSocket once connected

177 ws.send(

178 JSON.stringify({

179 type: "session.update",

180 session: {

181 type: "realtime",

182 instructions: "Be extra nice today!",

183 },

184 })

185 );

186});

187 

188// Listen for and parse server events

189ws.on("message", function incoming(message) {

190 console.log(JSON.parse(message.toString()));

191});

192```

193 

194 

195The WebSocket interface is perhaps the lowest-level interface available to interact with a Realtime model, where you will be responsible for both sending and processing Base64-encoded audio chunks over the socket connection.

196 

197To learn how to send and receive audio over Websockets, refer to the [Realtime conversations guide](https://developers.openai.com/api/docs/guides/realtime-conversations#handling-audio-with-websockets).

Details

398 398 

399## Custom voices399## Custom voices

400 400 

401Custom voices enable you to create a unique voice for your agent or application. These voices can be used for audio output with the [Text to Speech API](https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create), the [Realtime API](https://developers.openai.com/api/reference/resources/realtime), or the [Chat Completions API with audio output](https://developers.openai.com/api/docs/guides/audio).401Create an approved custom voice from a speaker's consent recording and matching

402 402audio sample. See [Custom voices](https://developers.openai.com/api/docs/guides/custom-voices) for eligibility,

403To create a custom voice, you’ll provide a short sample audio reference that the model will seek to replicate.403recording requirements, consent phrases, and API requests.

404 

405Custom voices are limited to eligible customers. Contact our [sales

406 team](https://openai.com/contact-sales/) to learn more. Once enabled for your

407 organization, you’ll have access to the

408 [Voices](https://platform.openai.com/audio/voices) tab under Audio.

409 404 

410#### Creating a voice405#### Creating a voice

411 406 

412Currently, voices must be created through an API request. See the API reference for the full set of API operations.407Follow [Create a custom voice](https://developers.openai.com/api/docs/guides/custom-voices#creating-a-voice).

413 

414Creating a voice requires two separate audio recordings:

415 

4161. **Consent recording** — this recording captures the voice actor providing consent to create a likeness of their voice. The actor must read one of the consent phrases provided below.

4172. **Sample recording** — the actual audio sample that the model will try to adhere to. The voice must match the consent recording.

418 

419**Tips for creating a high-quality voice**

420 

421The quality of your custom voice is highly dependent on the quality of the sample you provide. Optimizing the recording quality can make a big difference.

422 

423- Record in a quiet space with minimal echo.

424- Use a professional XLR microphone.

425- Stay about 7–8 inches from the mic with a pop filter in between, and keep that distance consistent.

426- The model copies exactly what you give it—tone, cadence, energy, pauses, habits—so record the exact voice you want. Be consistent in energy, style, and accent throughout.

427- Small variations in the audio sample can result in quality differences with the generated voice, it's worth trying multiple examples to find the best fit.

428 

429**Requirements and limitations**

430 

431- At most 20 voices can be created per organization.

432- The audio samples must be 30 seconds or less.

433- The audio samples must be one of the following types: `mpeg`, `wav`, `ogg`, `aac`, `flac`, `webm`, or `mp4`.

434 

435Refer to the Text-to-Speech Supplemental Agreement for additional terms of use.

436 

437**Creating a voice consent**

438 

439The consent audio recording must only include one of the following phrases. Any divergence from the script will lead to a failure.

440 

441| Language | Phrase |

442| -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |

443| `de` | Ich bin der Eigentümer dieser Stimme und bin damit einverstanden, dass OpenAI diese Stimme zur Erstellung eines synthetischen Stimmmodells verwendet. |

444| `en` | I am the owner of this voice and I consent to OpenAI using this voice to create a synthetic voice model. |

445| `es` | Soy el propietario de esta voz y doy mi consentimiento para que OpenAI la utilice para crear un modelo de voz sintética. |

446| `fr` | Je suis le propriétaire de cette voix et j'autorise OpenAI à utiliser cette voix pour créer un modèle de voix synthétique. |

447| `hi` | मैं इस आवाज का मालिक हूं और मैं सिंथेटिक आवाज मॉडल बनाने के लिए OpenAI को इस आवाज का उपयोग करने की सहमति देता हूं |

448| `id` | Saya adalah pemilik suara ini dan saya memberikan persetujuan kepada OpenAI untuk menggunakan suara ini guna membuat model suara sintetis. |

449| `it` | Sono il proprietario di questa voce e acconsento che OpenAI la utilizzi per creare un modello di voce sintetica. |

450| `ja` | 私はこの音声の所有者であり、OpenAIがこの音声を使用して音声合成 モデルを作成することを承認します。 |

451| `ko` | 나는 이 음성의 소유자이며 OpenAI가 이 음성을 사용하여 음성 합성 모델을 생성할 것을 허용합니다. |

452| `nl` | Ik ben de eigenaar van deze stem en ik geef OpenAI toestemming om deze stem te gebruiken om een synthetisch stemmodel te maken. |

453| `pl` | Jestem właścicielem tego głosu i wyrażam zgodę na wykorzystanie go przez OpenAI w celu utworzenia syntetycznego modelu głosu. |

454| `pt` | Eu sou o proprietário desta voz e autorizo o OpenAI a usá-la para criar um modelo de voz sintética. |

455| `ru` | Я являюсь владельцем этого голоса и даю согласие OpenAI на использование этого голоса для создания модели синтетического голоса. |

456| `uk` | Я є власником цього голосу і даю згоду OpenAI використовувати цей голос для створення синтетичної голосової моделі. |

457| `vi` | Tôi là chủ sở hữu giọng nói này và tôi đồng ý cho OpenAI sử dụng giọng nói này để tạo mô hình giọng nói tổng hợp. |

458| `zh` | 我是此声音的拥有者并授权OpenAI使用此声音创建语音合成模型 |

459 

460Then upload the recording via the API. A successful upload will return the consent recording ID that you’ll reference later. Note the consent can be used for multiple different voice creations if the same voice actor is making multiple attempts.

461 

462```bash

463curl https://api.openai.com/v1/audio/voice_consents \

464 -X POST \

465 -H "Authorization: Bearer $OPENAI_API_KEY" \

466 -F "name=test_consent" \

467 -F "language=en" \

468 -F "recording=@$HOME/tmp/voice_consent/consent_recording.wav;type=audio/x-wav"

469```

470 

471 

472**Creating a voice**

473 

474Next, you’ll create the actual voice by referencing the consent recording ID, and providing the voice sample.

475 

476```bash

477curl https://api.openai.com/v1/audio/voices \

478 -X POST \

479 -H "Authorization: Bearer $OPENAI_API_KEY" \

480 -F "name=test_voice" \

481 -F "audio_sample=@$HOME/tmp/voice_consent/audio_sample_recording.wav;type=audio/x-wav" \

482 -F "consent=cons_123abc"

483```

484 

485 

486If successful, the created voice will be listed under the [Audio tab](https://platform.openai.com/audio/voices).

487 408 

488#### Using a voice during speech generation409#### Using a voice during speech generation

489 410 

490Speech generation will work as usual. Simply specify the ID of the voice in the `voice` parameter when [creating speech](https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create), or when initiating a [realtime session](https://developers.openai.com/api/reference/resources/realtime/subresources/calls/methods/create#realtime_create_call-session-audio-output-voice).411Pass the created voice ID when generating speech. See the

491 412[speech generation examples](https://developers.openai.com/api/docs/guides/custom-voices#using-a-voice-during-speech-generation).

492**Text to speech example**

493 

494```bash

495curl https://api.openai.com/v1/audio/speech \

496 -X POST \

497 -H "Authorization: Bearer $OPENAI_API_KEY" \

498 -H "Content-Type: application/json" \

499 -d '{

500 "model": "gpt-4o-mini-tts",

501 "voice": {

502 "id": "voice_123abc"

503 },

504 "input": "Maple est le meilleur golden retriever du monde entier.",

505 "language": "fr",

506 "format": "wav"

507 }' \

508 --output sample.wav

509```

510 

511 

512**Realtime API example**

513 

514For Ruby, set `OPENAI_VOICE_ID` to your custom voice ID before running the example.

515 

516```javascript

517const sessionConfig = JSON.stringify({

518 session: {

519 type: "realtime",

520 model: "gpt-realtime-2",

521 audio: {

522 output: {

523 voice: { id: "voice_123abc" },

524 },

525 },

526 },

527});

528```

529 

530```ruby

531require "json"

532 

533session_config = JSON.generate(

534 session: {

535 type: "realtime",

536 model: "gpt-realtime-2",

537 audio: {output: {voice: {id: ENV.fetch("OPENAI_VOICE_ID")}}}

538 }

539)

540puts(session_config)

541```

542 

543 413 

544## Related guides414## Related guides

545 415 

Details

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Voice agents turn the same agent concepts into spoken, low-latency interactions. The key design choice is deciding whether the model should work directly with live audio or whether your application should explicitly chain speech-to-text, text reasoning, and text-to-speech.5Voice agents let users ask questions and complete tasks by speaking with your application. The key design choice is how speech connects to reasoning and tools: a continuous conversation with a separate backend, a single voice model, or a pipeline you control stage by stage.

6 6 

7## Choose the right architecture7## Choose the right architecture

8 8 

9| Architecture | Best for | Why |9| Architecture | Best for | Why choose it |

10| ----------------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------------------- |10| ---------------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |

11| Speech-to-speech with live audio sessions | Natural, low-latency conversations | The model handles live audio input and output directly |11| GPT-Live | Full-duplex conversations with a separate backend | Keep your existing text workflow and choose its backend independently while the conversation continues. |

12| Chained voice pipeline | Predictable workflows or extending an existing text agent | Your app keeps explicit control over transcription, text reasoning, and speech output |12| Realtime API | Speech, reasoning, and tool use in one session | Use one model to interpret audio, decide what to do, and respond in speech. |

13| Chained voice pipeline | Control over each speech and text stage | Inspect or transform intermediate text and replace each component independently. |

13 14 

14Voice workflows are an SDK-first surface. If you're migrating a related Agent Builder project, see [Migrate from Agent Builder](https://developers.openai.com/api/docs/guides/agent-builder/migrate-from-agent-builder) for the current transition path.

15 15 

16## Recommended starting points

17 16 

18The examples below are intentionally different architectures, not matching language tabs. The JavaScript and Python libraries expose different voice helpers today:

19 17 

20- In JavaScript, the fastest path to a browser-based voice assistant is a `RealtimeAgent` and `RealtimeSession`.

21- In Python, the simplest path to extending an existing text agent into voice is a chained `VoicePipeline`.

22 18 

19## Build a full-duplex voice agent

23 20 

21GPT-Live can listen and speak at the same time, a capability called **full duplex**. The live model handles the spoken interaction and delegates reasoning and tool use to a separate backend. Users can keep talking while backend work runs.

24 22 

23You can keep your existing text workflow, including its business logic and tools, and add GPT-Live as the voice interface. Your **delegation mode** determines who runs the backend work and supplies its conversation context:

25 24 

25- **Client delegation:** Connect your own agent or workflow, using the backend model and provider you choose. Your application runs the work and returns results to GPT-Live.

26- **Responses delegation:** Choose an OpenAI-hosted Responses model for backend reasoning and tool use. GPT-Live supplies conversation context and manages calls to that model; your application still runs custom functions.

26 27 

27## Build a speech-to-speech voice agent28In both modes, your application controls permissions and business records. Keep speaking behavior in the live model's prompt and business rules in the backend prompt.

28 

29Use the live audio API path when the interaction should feel conversational and immediate. This is the best starting point for voice agents that need barge-in, low first-audio latency, natural turn taking, and realtime tool use.

30 

31The usual browser flow is:

32 

331. Your application server creates an ephemeral client secret for the live audio session.

342. Your frontend creates a `RealtimeSession`.

353. The session connects over WebRTC in the browser or WebSocket on the server.

364. The agent handles audio turns, tools, interruptions, and handoffs inside that session.

37 29 

38Start a realtime voice session30Start with [Getting started with GPT-Live](https://developers.openai.com/api/docs/guides/live). See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for backend setup and [Prompting voice models](https://developers.openai.com/api/docs/guides/live-prompting) for speaking behavior.

39 31 

40```javascript

41import { RealtimeAgent, RealtimeSession } from "@openai/agents/realtime";

42 32 

43const agent = new RealtimeAgent({

44 name: "Assistant",

45 instructions: "You are a helpful voice assistant.",

46});

47 

48const session = new RealtimeSession(agent, {

49 model: "gpt-realtime-2.1",

50});

51 

52await session.connect({

53 apiKey: "ek_...(ephemeral key from your server)",

54});

55```

56 33 

57 34 

58From there, attach tools, handoffs, and guardrails to the `RealtimeAgent` the same way you would attach them to a text agent. Keep audio transport concerns in the session layer, and keep business logic in the agent definition.

59 35 

60Start with the transport docs when you need lower-level control:36## Build a speech-to-speech voice agent

61 37 

62- [Realtime and audio overview](https://developers.openai.com/api/docs/guides/realtime)38For the Realtime API, a `RealtimeAgent` and `RealtimeSession` provide a browser-first starting point. The session handles audio turns, tools, interruptions, and handoffs. The complete starter now lives in [Realtime API getting started](https://developers.openai.com/api/docs/guides/realtime#build-a-speech-to-speech-voice-agent).

63- [Live audio API with WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc)

64- [Live audio API with WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket)

65 39 

66## Build a chained voice workflow40## Build a chained voice workflow

67 41 

68Use the chained path when you want stronger control over intermediate text, existing text-agent reuse, or a simpler extension path from a non-voice workflow. In that design, your application explicitly manages:42Use the chained path when you want to inspect or transform text between speech recognition, your agent, and speech generation. Your application manages three stages:

69 

701. speech-to-text

712. the agent workflow itself

723. text-to-speech

73 43 

74This is often the better fit for support flows, approval-heavy flows, or cases where you want durable transcripts and deterministic logic between each stage.441. Speech-to-text

452. The agent workflow itself

463. Text-to-speech

75 47 

76Run a chained voice pipeline48Run a chained voice pipeline

77 49 


113 85 

114Use this path when each stage needs to be visible or replaceable. For example, you might store the transcript, run policy checks before the text agent responds, call internal systems, then generate speech only after the workflow reaches an approved answer.86Use this path when each stage needs to be visible or replaceable. For example, you might store the transcript, run policy checks before the text agent responds, call internal systems, then generate speech only after the workflow reaches an approved answer.

115 87 

88## Evaluate your voice agent

89 

90Test conversation quality and task outcomes separately. A natural-sounding response does not prove that a tool ran or that application state changed.

91 

921. Choose representative scenarios with expected outcomes, tool calls, and permissions.

932. Save the audio, events, tool results, and application state needed to verify each outcome. Distinguish a failed evaluation run from a valid run in which the agent fails the task.

943. Repeat scenarios and compare task completion, audible response latency, interruptions, and unwanted silence. Keep the caller, model configuration, tools, and transport consistent when comparing changes.

95 

96For GPT-Live, measure these dimensions independently:

97 

98- **Task and tool outcomes:** Check intent preservation, delegated work, tool arguments, permissions, and final application state. Verify that spoken confirmations match completed actions.

99- **Conversational timing:** Measure [audible response timing](#measure-latency), unwanted silence, overlap, and yielding to interruptions, including corrections while backend work runs.

100- **Speech and language:** Test input recognition across accents, background noise, language switches, names, and numbers. Assess output intelligibility and language choice separately from recognition.

101- **Session reliability:** Track connection failures, dropped audio, timeouts, and incomplete sessions separately from task scores.

102 

103Use **Crawl, Walk, and Run** to add complexity in stages:

104 

1051. **Crawl:** Use synthetic speech for controlled, single-turn requests. Keep the generated audio, application context, and expected outcome fixed for repeatable comparisons.

1062. **Walk:** Replay representative human recordings of single-turn requests to test how voices, microphones, pauses, and acoustic conditions affect behavior.

1073. **Run:** Use an independent simulated caller for continuous, multi-turn conversations. Test clarification, changing requirements, interruptions, and recovery while conversation and backend work overlap.

108 

109Complement automated scores with human listening to assess pronunciation, naturalness, and whether the conversation feels appropriately paced.

110 

111For a GPT-Live evaluation harness, see the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation).

112 

113For a Realtime evaluation harness and worked examples, use the [Realtime evaluation guide in the OpenAI Cookbook](https://developers.openai.com/cookbook/examples/realtime_eval_guide). The Cookbook owns the runnable evaluation recipes; this page provides the shared testing checklist.

114 

115### Measure latency

116 

117Define an observed start and end event for every latency metric. Time to first

118audible response, time to delegation, interruption yield, backend completion,

119and verified task completion measure different boundaries. Use one monotonic

120timeline and report the eligible population, median, and tail latency. Do not

121substitute a backend-only timer for end-to-end response time.

122 

123Keep the caller, recording, backend model, prompt, transport, audio cadence, and

124grader fixed when comparing frontend models.

125 

126For GPT-Live, record the stages your application can observe: delegation receipt,

127backend request start, first useful result, tool start and end, result submission,

128audio arrival, and client playback. Client delegation gives your application

129direct visibility into its backend requests; Responses delegation exposes nested

130response events and the custom tools your application runs.

131 

132Use the intervals to locate delays in connection setup, model work, tools,

133application buffering, and playback. Measure the first useful spoken answer

134separately from an acknowledgment such as “I'm checking.” An earlier

135acknowledgment does not show that the requested result arrived sooner.

136 

137Change one factor at a time and repeat the same scenarios. Compare median and

138tail time to useful spoken responses alongside task success, tool correctness,

139and interruptions. See [Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency)

140for implementation guidance.

141 

116## Voice agents still use the same core agent building blocks142## Voice agents still use the same core agent building blocks

117 143 

118The voice surface changes the transport and audio loop, but the core workflow decisions are the same:144The voice surface changes the transport and audio loop, but the core workflow decisions are the same:


127 153 

128## Next steps154## Next steps

129 155 

130[Realtime and audio overview156[Audio and voice overview

131 157 

132 158 

133 159 

134 Choose the right realtime or audio guide for your use case.](https://developers.openai.com/api/docs/guides/realtime)160 Choose the right realtime or audio guide for your use case.](https://developers.openai.com/api/docs/guides/audio)

135 161 

136[Managing conversations162[Managing conversations

137 163 


143 169 

144 170 

145 171 

146 Connect browser and mobile audio directly to a Realtime session.](https://developers.openai.com/api/docs/guides/realtime-webrtc)172 Connect browser and mobile audio directly to a Realtime session.](https://developers.openai.com/api/docs/guides/voice-webrtc)

147 173 

148[Realtime prompting guide174[Realtime prompting guide

149 175 

150 176 

151 177 

152 Tune reasoning, preambles, tools, entity capture, and voice behavior.](https://developers.openai.com/api/docs/guides/realtime-models-prompting)178 Tune reasoning, preambles, tools, entity capture, and voice behavior.](https://developers.openai.com/api/docs/guides/voice-prompting)

guides/voice-latency-cost.md +201 −1 renamed

Details

Previously: guides/realtime-costs.md

1# Managing costs1# Cost optimization

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose your API to understand how usage is measured and find ways to manage

6costs for your voice application.

7 

8 

9 

10## GPT-Live usage and costs

11 

12GPT-Live separates the voice conversation from the backend that reasons and

13runs tools. Estimate these two costs separately: the voice session depends on

14duration, while backend costs depend on the models and tools you use.

15 

16### Voice session costs

17 

18GPT-Live voice sessions are billed per second at the current [model rate](https://developers.openai.com/api/docs/models/gpt-live-1). Session duration is not rounded up to the next whole minute.

19 

20Active session time includes time when the user speaks, the assistant speaks, both are silent,

21or the backend is working.

22 

23For estimates, count the active session from start through closure. Use the

24duration reported by the API instead of timing only the audio you play. Muting

25microphone input does not close the session. When the conversation is finished,

26close the session and collect its final usage.

27 

28See [API pricing](https://developers.openai.com/api/docs/pricing) for backend model and tool prices.

29 

30### WebRTC initialization charges

31 

32A `POST /v1/live/sessions` request to create a WebRTC session bills 15 seconds of voice duration while the session initializes. That amount is credited against duration charges once the session starts running. Don't add another 15 seconds to the running session's duration when estimating its cost.

33 

34For example, the 90-second session below already includes the 15 seconds billed at initialization. It is not billed as 105 seconds. Account for session-creation charges when evaluating reconnects or applications that create sessions before the user is ready to speak.

35 

36### Backend costs

37 

38Backend calls are billed separately from the voice session, just as they are in

39applications without voice. Include model input and output tokens, cached input

40where supported, and any applicable image or tool charges. If your application

41calls other services, include their costs in your estimate too.

42 

43You can optimize this work separately from the voice frontend. Use the general

44[cost optimization guide](https://developers.openai.com/api/docs/guides/cost-optimization) to reduce requests

45and token usage. Use [prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) for eligible

46backend models by keeping reusable instructions, tool definitions, and other

47stable content at the beginning of the prompt.

48 

49Backend choices can also change the length of the conversation. Compare the

50combined cost when an optimization makes the user wait longer or changes how

51reliably the assistant completes the task.

52 

53### Estimate conversation costs

54 

55For a conversation with one voice session:

56 

57**Total cost = (billable voice seconds ÷ 60 × voice rate per minute) + backend costs**

58 

59For example, at an illustrative voice rate of $0.05 per minute, a 90-second voice session costs $0.075. If the backend model and tool costs total $0.02, the conversation costs $0.095:

60 

61| Component | Calculation | Cost |

62| ---------------------- | -------------------------- | ---------- |

63| Voice session | 90 seconds ÷ 60 × $0.05 | $0.075 |

64| Backend work | Total model and tool costs | $0.02 |

65| **Conversation total** | **$0.075 + $0.02** | **$0.095** |

66 

67The rates and backend cost above are examples; use the current voice rate, your measured backend usage, and the

68applicable model and tool rates. If the task spans multiple voice sessions, add

69their durations and include backend work performed between sessions.

70 

71### Optimization strategies

72 

73Focus on helping the user complete the task with less unnecessary conversation

74and waiting. Keep the confirmations and checks the task requires.

75 

76#### Provide relevant context before the session

77 

78Gather information your application already has permission to use before

79starting the voice session. For example, an assistant helping with an order can

80start with the order number and current status, so the user does not need to

81repeat them or wait for another lookup.

82 

83Keep this context current and focused on the task. Give the voice model the

84information it needs for the conversation; keep detailed records and workflows

85in the backend. See [session configuration](https://developers.openai.com/api/docs/guides/live-conversations#session-configuration)

86and [delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation).

87 

88#### Reduce time spent waiting for tools

89 

90Shorter waits can improve the user experience and reduce voice-session costs.

91For example, suppose your backend uses `gpt-5.6-luna` with

92[Fast mode](https://developers.openai.com/api/docs/guides/fast-mode) and runs independent tool calls in

93parallel. If these optimizations help the user finish and close the voice

94session one minute sooner, you save $0.05 in voice charges. The total cost

95falls if the additional backend cost is less than that saving.

96 

97You can also [start a speculative lookup from transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments)

98before a delegation event arrives. Include unused speculative work in your

99backend cost measurements.

100 

101See [Reduce backend latency](https://developers.openai.com/api/docs/guides/live-delegation#reduce-backend-latency)

102for model, connection, streaming, and tool optimizations. Validate useful spoken

103response time and task success with [voice agent evaluations](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation).

104 

105#### Close the session during long tasks

106 

107The voice frontend and your application-managed backend can run independently.

108With client delegation, your backend worker can keep running while the voice

109session is open or closed. Save the task state and conversation context before

110[closing the voice session](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close).

111 

112For an ambient agent, close the voice session while the backend handles a

113long-running task, such as coding in goal mode. Offer a button labeled

114**Resume conversation** to start a new voice session when the user returns, or use a backend

115completion event to start a new session and notify the user that the result is

116ready.

117 

118Restore the conversation by starting a new session with saved context and the

119verified task result in `input`. For example, send this startup event over a

120[new WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live):

121 

122```json

123{

124 "type": "session.start",

125 "session": {

126 "model": "gpt-live-1",

127 "instructions": "Help the user review completed work and delegate follow-up tasks.",

128 "input": [

129 {

130 "type": "message",

131 "role": "developer",

132 "content": [

133 {

134 "type": "input_text",

135 "text": "Saved task: add CSV export. Result: code is ready for review."

136 }

137 ]

138 }

139 ],

140 "delegation": { "type": "client" }

141 }

142}

143```

144 

145Wait for `session.started` before streaming audio. See

146[seed a session with prior conversation](https://developers.openai.com/api/docs/guides/live-conversations#seed-a-session-with-prior-conversation)

147for the supported history format.

148 

149If the earlier session was stored with `store: true`, you can also [fork that session](https://developers.openai.com/api/docs/guides/live-conversations#store-and-fork-a-session). Keep the verified backend task state in your application whichever approach you use.

150 

151Closing saves $0.05 per minute of idle voice time; compare that saving with

152reconnection costs and the interruption to the user's experience.

153 

154#### Choose the right backend model

155 

156Start with models that meet the task's accuracy and reliability requirements.

157Then compare total conversation cost, including voice duration, model usage,

158tool calls, and retries. The [model selection guide](https://developers.openai.com/api/docs/guides/model-selection)

159describes how to balance these tradeoffs.

160 

161A larger backend model can cost less overall if it completes the task faster

162and the voice-session savings exceed its additional token costs. A cheaper

163model can cost more overall if it takes longer, repeats tool calls, or fails

164the task.

165 

166Compare cost per successful task alongside completion rate and time to completion. Include failed attempts and retries in the total so a cheaper configuration does not look better because it completes less work. Use the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) when planning your comparison.

167 

168### Monitor actual usage

169 

170Record voice duration and backend usage separately for each session. GPT-Live

171reports cumulative voice duration in seconds:

172 

173```json

174{

175 "type": "session.usage.updated",

176 "event_id": "event_usage_1",

177 "usage": { "seconds": 12 },

178 "context_window": { "usage_ratio": 0.42 }

179}

180```

181 

182Each update replaces the previous duration snapshot. Do not sum the snapshots.

183After sending `session.close`, keep receiving events until `session.closed` and

184record its final `usage.seconds` once. Follow the

185[graceful-close procedure](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close)

186so your application can collect final usage before disconnecting.

187 

188For Responses delegation, read the backend response's `usage` from nested

189`response.completed` events delivered through `response.event`. Count each

190backend response once, using its response ID, and retain the input, output, and

191cached-token details needed to apply that model's rates. For backend work your

192application runs independently, collect usage from those requests too.

193 

194Compare estimated and actual totals across representative conversations. Keep

195evaluation-only model calls separate from application usage, and review cost

196together with task success.

197 

198

199 

200

201 

202 

203## Realtime API costs

204 

5This document describes how Realtime API billing works and offers strategies for optimizing costs. Voice-agent sessions accrue input and output tokens across text, audio, and image modalities. Streaming translation and streaming transcription sessions are billed by audio duration. Prices vary per model, with prices listed on the model pages (for example, [`gpt-realtime-2`](https://developers.openai.com/api/docs/models/gpt-realtime-2), [`gpt-realtime-translate`](https://developers.openai.com/api/docs/models/gpt-realtime-translate), [`gpt-realtime-whisper`](https://developers.openai.com/api/docs/models/gpt-realtime-whisper), and [`gpt-realtime`](https://developers.openai.com/api/docs/models/gpt-realtime)).205This document describes how Realtime API billing works and offers strategies for optimizing costs. Voice-agent sessions accrue input and output tokens across text, audio, and image modalities. Streaming translation and streaming transcription sessions are billed by audio duration. Prices vary per model, with prices listed on the model pages (for example, [`gpt-realtime-2`](https://developers.openai.com/api/docs/models/gpt-realtime-2), [`gpt-realtime-translate`](https://developers.openai.com/api/docs/models/gpt-realtime-translate), [`gpt-realtime-whisper`](https://developers.openai.com/api/docs/models/gpt-realtime-whisper), and [`gpt-realtime`](https://developers.openai.com/api/docs/models/gpt-realtime)).

6 206 

7Conversational Realtime API sessions are a series of _turns_, where the user adds input that triggers a _Response_ to produce the model output. The server maintains a _Conversation_, which is a list of _Items_ that form the input for the next turn. When a Response is returned, the output is automatically added to the Conversation.207Conversational Realtime API sessions are a series of _turns_, where the user adds input that triggers a _Response_ to produce the model output. The server maintains a _Conversation_, which is a list of _Items_ that form the input for the next turn. When a Response is returned, the output is automatically added to the Conversation.

guides/voice-prompting.md +19 −4 renamed

Details

Previously: guides/realtime-models-prompting.md

1# Using realtime models1# Prompting Realtime models

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose the Realtime model you are building with. For GPT-Live, use [Prompting GPT-Live](https://developers.openai.com/api/docs/guides/live-prompting).

6 

7 

8 

9 

10 

11 

12 

5`gpt-realtime-2` is our state-of-the-art reasoning voice model for low-latency speech-to-speech applications. It can think before it speaks, follow instructions more reliably, use a larger context window, and call tools with greater precision than earlier realtime models.13`gpt-realtime-2` is our state-of-the-art reasoning voice model for low-latency speech-to-speech applications. It can think before it speaks, follow instructions more reliably, use a larger context window, and call tools with greater precision than earlier realtime models.

6 14 

7To take advantage of these gains, design prompts with more intent. Explicitly define the assistant's responsibilities, decision points, tool-calling behavior, and guardrails: what it should do, when it should do it, and what it should avoid.15To take advantage of these gains, design prompts with more intent. Explicitly define the assistant's responsibilities, decision points, tool-calling behavior, and guardrails: what it should do, when it should do it, and what it should avoid.


2081 2087 

2082## Next steps2088## Next steps

2083 2089 

2090For GPT-Live:

2091 

2092- Review [delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for the backend prompt and application-owned context.

2093- Connect with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) or [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live). See [Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip?api=live) for phone integrations.

2094- [Evaluate voice agents](https://developers.openai.com/api/docs/guides/voice-agents#evaluate-your-voice-agent) across conversation quality and verified task outcomes.

2095 

2096For Realtime:

2097 

2084- Review the earlier [Realtime prompting guide](https://developers.openai.com/cookbook/examples/realtime_prompting_guide) for more `gpt-realtime-1.5` examples.2098- Review the earlier [Realtime prompting guide](https://developers.openai.com/cookbook/examples/realtime_prompting_guide) for more `gpt-realtime-1.5` examples.

2085- Review the [Realtime eval guide](https://developers.openai.com/cookbook/examples/realtime_eval_guide) to test representative voice-agent behavior.2099- Review the [Realtime eval guide](https://developers.openai.com/cookbook/examples/realtime_eval_guide) to test representative voice-agent behavior.

2086- Learn how to connect with [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc), [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket), or [SIP](https://developers.openai.com/api/docs/guides/realtime-sip).2100- Connect with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime), [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime), or [SIP](https://developers.openai.com/api/docs/guides/voice-sip?api=realtime).

2087- Learn the [Realtime conversation lifecycle](https://developers.openai.com/api/docs/guides/realtime-conversations).2101- Learn the [Realtime conversation lifecycle](https://developers.openai.com/api/docs/guides/realtime-conversations) and review [Realtime costs](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=realtime).

2088- Review [Realtime costs](https://developers.openai.com/api/docs/guides/realtime-costs).

Details

1# Server-side controls

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.

6 

7 

8 

9## Control a GPT-Live session from your server

10 

11Attach your application server to an existing GPT-Live WebRTC or SIP session when the server needs to receive conversation events, execute private tools, or update the conversation. This second connection is called a **sideband WebSocket**. Both connections share one session while WebRTC or SIP carries the primary audio.

12 

13The sideband carries events and commands. Your application supplies the tool execution, authorization checks, and business rules. Keep API keys and tool credentials on your server.

14 

15### Decide whether you need a sideband

16 

17For browser applications, use the [WebRTC data channel](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live) for captions and local UI updates. Use a sideband when transcript processing runs on your server, such as guardrail checks, sentiment analysis, or speculative tool calls. Your server can receive events and steer the same session directly while browser audio stays on WebRTC. See [React to transcript fragments](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments) for examples.

18 

19If your backend already owns the primary [WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live), it already receives the session's events and can send commands.

20 

21[Responses delegation](https://developers.openai.com/api/docs/guides/live-delegation) also works without a sideband. The browser can forward function-call events from its data channel to an authenticated backend for execution. OpenAI-hosted tools run through the delegated backend without an application tool executor.

22 

23### Attach to the existing session

24 

251. Save the ID of the session your backend will control. For WebRTC, use `session.id` from the JSON response to `POST /v1/live/sessions`. For SIP, [accept the incoming call](https://developers.openai.com/api/docs/guides/voice-sip?api=live#accept-or-reject-the-call) first, then use `data.session_id` from its webhook. Keep the ID alongside the application's user and conversation record.

262. Open a WebSocket from your server at the following URL, substituting the saved ID unchanged. Authenticate with `Authorization: Bearer $OPENAI_API_KEY` using the project authentication that created or accepted the session. Include the same connection headers required when creating the session.

27 

28```text

29 wss://api.openai.com/v1/live/sessions/{session_id}/attach

30```

31 

323. Receive events and send commands on the attached socket. The session is already running; do not send `session.start` again.

33 

34Treat the session ID as an opaque value. Preserve its prefix and use it only for the session to which your application has authorized access. Read the ID from the Live JSON response, rather than a Realtime `Location` header or `call_id` URL parameter.

35 

36### Observe events and send commands

37 

38| Task | Events or commands |

39| ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

40| Follow the conversation | Receive user and assistant transcript deltas, delegation events, and nested Responses events. |

41| Update backend configuration | Use `session.update` to change supported settings within the existing delegation mode. Startup settings such as the frontend model and audio configuration stay fixed. |

42| Provide context | Use `session.instructions.append` for instructions, `session.thinking.append` for quiet context, and `session.commentary.append` for speakable updates. |

43| Return tool results | With Responses delegation, send `response.item.create`, then `response.create` to continue backend work. |

44| Control microphone input | Use `session.input_audio.mute` and `session.input_audio.unmute`. Muting input does not stop the assistant's output. |

45| Finish the session | Send `session.close` and receive `session.closed` before disconnecting. |

46 

47Commands follow the same validation and delegation rules as on the primary connection. For context appends, use `delegation_id: null` for general session context; a non-null ID must identify an existing client delegation. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for configuration, function execution, and the context append examples.

48 

49For browser sessions, keep microphone input and speaker output on the negotiated WebRTC media track. Use the sideband for conversation events and control. A transcript event or command acknowledgment does not prove that audio has played or that the user has heard it.

50 

51### Receive reflected audio

52 

53A sideband also receives copies of subsequent input and output audio while the primary connection carries the live media:

54 

55| Event | Audio field | Timing |

56| ---------------------------- | ----------- | ---------------------------------------------------------------------------- |

57| `session.input_audio.append` | `audio` | No timestamps. |

58| `session.output_audio.delta` | `delta` | `start_ms` and `end_ms` describe the output's range on the session timeline. |

59 

60Both payloads are base64-encoded raw mono PCM16LE at 24 kHz, regardless of the primary transport's audio format. Neither event has an `event_id`. Reflected input contains received audio before input muting; it does not confirm that the model consumed those samples. Reflected output ranges can have gaps for dropped frames and do not indicate when the caller heard the audio.

61 

62These are server events, not permission to send audio through the sideband. Send microphone audio through the primary transport; do not send `session.input_audio.append` on the attached socket.

63 

64### Assign one owner for each action

65 

66Choose whether the browser or backend handles each action. If both connections receive a function-call event, execute the function once. Apply the same ownership rule to context updates and requests to continue backend work.

67 

68Store transcripts and tool state in your application. Attach early if the backend needs to observe the conversation from the start, and retain any history collected before attachment. Do not rely on attachment to reconstruct earlier transcripts or tool results.

69 

70A sideband does not itself make session events private from the browser. Keep sensitive tool credentials and authorization decisions in your backend, and return only the context needed for the conversation.

71 

72 

73 

74 

75 

76## Apply conversation guardrails

77 

78Use your server's connection to monitor the conversation, check requests against your application's policies, and intervene when a check triggers. A sideband gives your server access to session events and commands; your application runs the checks and enforces their results. The same workflow applies when your server already owns the primary WebSocket connection.

79 

80### Run checks alongside the conversation

81 

82Guardrails are one use of [processing transcript fragments as they arrive](https://developers.openai.com/api/docs/guides/live-delegation#react-to-transcript-fragments). The same stream can start a speculative lookup or update the UI alongside these checks.

83 

841. **Monitor transcripts.** Accumulate `session.input_transcript.delta` fragments to check user requests for jailbreak attempts, sensitive information, or policy violations. Use `session.output_transcript.delta` to check assistant speech for unsupported claims or responses outside your application's scope. Keep each check associated with the transcript and application request it evaluated.

852. **Run checks concurrently.** A fast, lightweight model can evaluate requests while the conversation continues. Return a small structured result, such as `{"triggered": true}`, that your application can act on. Keep actions that require approval blocked until their checks pass; a timeout or failed check is not approval.

863. **Block affected actions.** When a check triggers, mark the request as blocked in application state. Check that state before executing a tool or committing a change, including work already queued. A spoken refusal does not prevent a tool from running.

874. **Stop related work.** Cancel application-owned jobs where your backend supports cancellation, and discard late results from blocked or superseded requests. With Responses delegation, stop executing affected custom functions and do not send `response.create` to continue blocked work. This does not cancel an already-running hosted response or stop frontend speech.

885. **Record and redirect.** Log the decision with the affected request and delegation IDs, then send a corrective instruction. An event name such as `guardrail.triggered` belongs to your application's telemetry; it is not a GPT-Live API event.

89 

90See [Transcript deltas](https://developers.openai.com/api/docs/guides/live-conversations#transcript-deltas) for collecting fragments and [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation#keep-updates-accurate-and-useful) for keeping backend results aligned with the current task.

91 

92### Redirect the conversation

93 

94Use `session.instructions.append` for guardrail steering. It can interrupt speech in progress and apply a new instruction. For example, after your application blocks a request, send:

95 

96```javascript

97/**

98 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

99 */

100export function sendUpdate(connection) {

101 connection.send({

102 type: "session.instructions.append",

103 event_id: "guardrail_block_17",

104 delegation_id: null,

105 content:

106 "Stop speaking immediately. Do not continue or act on the last request. Refuse briefly, then wait.",

107 });

108}

109```

110 

111```python

112from openai.resources.live.live import AsyncLiveConnection

113from openai.resources.live.sideband import AsyncSidebandConnection

114 

115 

116async def send_update(

117 connection: AsyncLiveConnection | AsyncSidebandConnection,

118) -> None:

119 await connection.session.instructions.append(

120 event_id="guardrail_block_17",

121 delegation_id=None,

122 content=(

123 "Stop speaking immediately. Do not continue or act on the last request. "

124 "Refuse briefly, then wait."

125 ),

126 )

127```

128 

129 

130Keep the instruction application-authored. Do not copy untrusted user text into it as an instruction. Use `delegation_id: null` for this session-wide correction, and keep `content` within 500 tokens.

131 

132Match `session.instructions.appended` to your command through `client_event_id`. The acknowledgment arrives after estimated context injection; it does not prove that the assistant stopped speaking or that queued audio stopped playing. Corrective instructions cannot retract audio the user has already heard.

133 

134For disclosures that request specific spoken wording, also use instructions. See [Deliver a disclosure](https://developers.openai.com/api/docs/guides/live-conversations#deliver-a-disclosure) for an example and playback considerations.

135 

136### Control playback when needed

137 

138Test corrective instructions and action blocking first. If your application also needs to block model audio, control output at the client or media relay: temporarily mute or drop the output, discard locally queued audio, send the corrective instruction, and resume playback according to your application's recovery policy. Clear stale audio before resuming. A sideband alone does not control the media path, and an instruction acknowledgment is not a signal to resume playback.

139 

140`session.input_audio.mute` controls the caller's microphone input. It does not mute model output or cancel delegated work.

141 

142GPT-Live streams transcript fragments while speaking. If a check must finish before the user hears the audio, your application needs to buffer and approve audio before playback. This adds latency. Suppressed audio can also leave the model's conversation context ahead of what the user heard, so test how the conversation resumes.

143 

144### Test the intervention

145 

146Test allowed and blocked requests, false positives, slow or failed checks, a trigger during speech, a trigger while a tool is running, and late results from canceled work. Verify action blocking, application state, corrective speech, and actual playback separately. If you control output, include queued audio and recovery in the test. Use the [voice agent evaluation Cookbook](https://developers.openai.com/cookbook/examples/audio/voice_agent_evaluation) to compare task success and spoken response time.

147 

148## Finish cleanly

149 

150Keep receiving events while the backend owns tool execution or final usage collection. Register the `session.closed` handler before sending `session.close`, and keep the WebRTC connection, data channel, and sideband open while pending work drains. Save the final session usage and any backend usage received in Responses events before cleanup. If the connection fails before the final event arrives, record finalization as incomplete. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for the close sequence.

151 

152

153 

154

155 

156 

157The Realtime API allows clients to connect directly to the API server via WebRTC or SIP. However, you'll most likely want tool use and other business logic to reside on your application server to keep this logic private and client-agnostic.

158 

159Keep tool use, business logic, and other details secure on the server side by connecting over a “sideband” control channel. We now have sideband options for both SIP and WebRTC connections.

160 

161A sideband connection means there are two active connections to the same Realtime session: one from the user's client and one from your application server. The server connection can be used to monitor the session, update instructions, and respond to tool calls.

162 

163## With WebRTC

164 

1651. When [establishing a peer connection](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime) you fetch and receive an SDP response from the Realtime API to configure the connection. If you used the sample code from the WebRTC guide, that looks something like this:

166 

167```javascript

168const baseUrl = "https://api.openai.com/v1/realtime/calls";

169const sdpResponse = await fetch(baseUrl, {

170 method: "POST",

171 body: offer.sdp,

172 headers: {

173 Authorization: `Bearer ${EPHEMERAL_KEY}`,

174 "Content-Type": "application/sdp",

175 },

176});

177```

178 

179 

1802. The fetch response will contain a `Location` header that has a unique call ID that can be used on the server to establish a WebSocket connection to that same Realtime session.

181 

182```javascript

183// Location: /v1/realtime/calls/rtc_123456

184const location = sdpResponse.headers.get("Location");

185const callId = location?.split("/").pop();

186console.log(callId);

187```

188 

189 

1903. On a server, you can then [listen for events and configure the session](https://developers.openai.com/api/docs/guides/realtime-conversations) just as you would from a typical Realtime API WebSocket connection, using that call ID with the URL

191 `wss://api.openai.com/v1/realtime?call_id=rtc_xxxxx`, as shown below:

192 

193```javascript

194import WebSocket from "ws";

195const callId = "rtc_u1_9c6574da8b8a41a18da9308f4ad974ce";

196 

197// Connect to a WebSocket for the in-progress call

198const url = "wss://api.openai.com/v1/realtime?call_id=" + callId;

199const ws = new WebSocket(url, {

200 headers: {

201 Authorization: "Bearer " + process.env.OPENAI_API_KEY,

202 },

203});

204 

205ws.on("open", function open() {

206 console.log("Connected to server.");

207 

208 // Send client events over the WebSocket once connected

209 ws.send(

210 JSON.stringify({

211 type: "session.update",

212 session: {

213 type: "realtime",

214 instructions: "Be extra nice today!",

215 },

216 })

217 );

218});

219 

220// Listen for and parse server events

221ws.on("message", function incoming(message) {

222 console.log(JSON.parse(message.toString()));

223});

224```

225 

226 

227In this way, you are able to add tools, monitor sessions, and carry out business logic on the server instead of needing to configure those actions on the client.

228 

229## With SIP

230 

2311. A user connects to OpenAI via phone over SIP.

2322. OpenAI sends a webhook to your application’s server webhook URL, notifying your app of the state of the session. The webhook will look something like:

233 

234```json

235POST https://my_website.com/webhook_endpoint

236user-agent: OpenAI/1.0 (+https://platform.openai.com/docs/webhooks)

237content-type: application/json

238webhook-id: wh_685342e6c53c8190a1be43f081506c52 # unique id for idempotency

239webhook-timestamp: 1750287078 # timestamp of delivery attempt

240webhook-signature: v1,K5oZfzN95Z9UVu1EsfQmfVNQhnkZ2pj9o9NDN/H/pI4= # signature to verify authenticity from OpenAI

241 

242{

243 "object": "event",

244 "id": "evt_685343a1381c819085d44c354e1b330e",

245 "type": "realtime.call.incoming",

246 "created_at": 1750287018, // Unix timestamp

247 "data": {

248 "call_id": "some_unique_id",

249 "sip_headers": [

250 { "name": "From", "value": "sip:+142555512112@sip.example.com" },

251 { "name": "To", "value": "sip:+18005551212@sip.example.com" },

252 { "name": "Call-ID", "value": "03782086-4ce9-44bf-8b0d-4e303d2cc590"}

253 ]

254 }

255}

256 

257```

258 

2593. The application server opens a WebSocket connection to the Realtime API using the `call_id` value provided in the webhook, via a URL like this: `wss://api.openai.com/v1/realtime?call_id={callId}`. The WebSocket connection will live for the life of the SIP call.

260 

261The WebSocket connection can then be used to send and receive events to control the call, just as you would if the session was initiated with a WebSocket connection. This includes monitoring the call, updating instructions dynamically, and responding to tool calls.

guides/voice-sip.md +109 −5 renamed

Details

Previously: guides/realtime-sip.md

1# Realtime API with SIP1# Telephony and SIP

2 2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.

6 

7 

8 

9## Choose a telephony connection

10 

11A phone call can reach GPT-Live through a SIP trunk or through an application that relays audio. Choose the path that fits your existing phone system and where your application needs to process audio.

12 

13| Connection | Audio and application responsibilities |

14| ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |

15| Direct SIP | The provider exchanges call audio with OpenAI. Your application handles webhooks, session configuration, call decisions, and business logic. |

16| Server audio bridge | Your application relays provider or room audio to GPT-Live over WebSocket. It manages both connections, event translation, playback, and call lifecycle. |

17 

18A provider's connection to your application and your application's connection to OpenAI are separate. For example, a caller can join a room through SIP while an agent in that room connects to GPT-Live over WebSocket.

19 

20Using Twilio, Telnyx, LiveKit, or Daily/Pipecat? See [GPT-Live partner integrations](https://developers.openai.com/api/docs/guides/live-partner-integrations) for provider-specific guides.

21 

22### Direct SIP

23 

24Direct SIP keeps call audio on the provider-to-OpenAI media path. SIP signaling uses TLS, and GPT-Live requires SRTP for call audio. Your backend still owns the incoming-call decision, session configuration, authorization, and business logic.

25 

26Use a [sideband connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) when your backend needs to receive session events or send commands. It attaches to the existing conversation while SIP carries the audio. Assign one handler to each action so that duplicate webhook deliveries or events observed on multiple connections don't execute tools twice.

27 

28Keep SIP routing and provider configuration together with the integration that uses them. Realtime webhook events, call identifiers, and acceptance payloads belong to the Realtime API; use the GPT-Live contract for a Live session.

29 

30### Handle the call lifecycle

31 

32Confirm that GPT-Live SIP support is enabled for your project and that your

33 provider's SIP trunk is routed to that project before using this flow. The

34 Realtime webhook and acceptance payloads in the other tab are a different API

35 contract.

36 

37#### Receive the incoming call

38 

39Configure your project's [webhook endpoint](https://developers.openai.com/api/docs/guides/webhooks) for `live.transport.incoming`. Verify the webhook signature and deduplicate deliveries before making a call decision. A delivery acknowledgment does not accept the call.

40 

41The webhook identifies a SIP call with `data.type: "sip"` and provides `data.session_id`. Use that session ID unchanged for every Live call action. Treat `data.sip_headers` as untrusted caller metadata, not authorization.

42 

43Existing integrations may still receive the deprecated `live.call.incoming` event, which has no `data.type`. During migration, handle both names and retain the old subscription until legacy deliveries and retries have drained. The same pending call can also emit a Realtime webhook; assign one handler to the accept/reject decision rather than accepting through both APIs.

44 

45#### Accept or reject the call

46 

47Apply your application's authorization and routing rules. To [accept the call](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/accept), send an authenticated `POST /v1/live/sessions/{session_id}/accept` request with a top-level `session` object:

48 

49```json

50{

51 "session": {

52 "type": "live",

53 "model": "gpt-live-1",

54 "instructions": "You are answering an inbound support call.",

55 "audio": { "output": { "voice": "marin" } },

56 "delegation": { "type": "client" }

57 }

58}

59```

60 

61Use `Authorization: Bearer $OPENAI_API_KEY` from your trusted backend for call-control requests. Choose the voice and delegation mode at acceptance. SIP negotiates the audio format, so omit `audio.format`. The example selects client delegation; your backend must handle delegated work. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for client and Responses configurations.

62 

63A successful acceptance returns `200 OK` with an empty body after session initialization. Handle HTTP errors before treating the call as accepted.

64 

65To [reject the call](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/reject), send `POST /v1/live/sessions/{session_id}/reject` with a SIP status, such as `{ "status_code": 486 }` for busy. The status must be an integer from 300 through 699. The first accept or reject decision wins; a later competing decision returns `decision_already_made`.

66 

67#### Attach your backend

68 

69After acceptance, connect a [sideband WebSocket](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) at `wss://api.openai.com/v1/live/sessions/{session_id}/attach`. Use the accepted session ID and the same project authentication and connection headers. Do not send `session.start` again.

70 

71SIP carries the call audio. Use the sideband for transcripts, delegation, tools, commands, and reflected audio. Choose one owner for each side effect, even if multiple connections observe an event.

72 

73#### Observe keypad events

74 

75The sideband receives `transport.dtmf.received` when the caller presses a key and `transport.dtmf.send` after a hosted tool successfully sends a tone. The event's `event` field contains one of `0`–`9`, `*`, `#`, or `A`–`D`.

76 

77These are observer notifications, not client commands. Do not send `transport.dtmf.send` to request a tone, or assume the browser data channel receives keypad events.

78 

79#### Transfer or end the call

80 

81To [transfer the call](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/refer), send `POST /v1/live/sessions/{session_id}/refer` with `{ "target_uri": "sip:agent@example.com" }` for your destination. To [hang up](https://developers.openai.com/api/reference/resources/live/subresources/sessions/methods/hangup), send `POST /v1/live/sessions/{session_id}/hangup` with no request body. Both return `200 OK` with an empty body on success.

82 

83Keep your sideband open for final events and usage before releasing application resources. A successful hangup request or an unexpected disconnect is not a substitute for `session.closed`. See [Usage and graceful close](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for finalization and close reasons.

84 

85This flow accepts inbound calls. Creating an outbound SIP call through `POST /v1/live/sessions` is not supported; use the relevant [partner integration](https://developers.openai.com/api/docs/guides/live-partner-integrations) for provider-owned outbound calling.

86 

87### Server audio bridges

88 

89Use the [GPT-Live WebSocket connection](https://developers.openai.com/api/docs/guides/voice-websockets?api=live) when your application receives an audio stream from a phone provider or an agent framework. The application authenticates both connections, translates their event envelopes, and relays audio in both directions.

90 

91GPT-Live supports raw G.711 μ-law and A-law audio at 8 kHz over WebSocket. When the provider stream uses the same codec, sample rate, and channel count, your application can forward the raw audio bytes without converting them to PCM. Preserve audio order and use the message format required by each connection. Matching audio formats don't make the two event protocols interchangeable.

92 

93The bridge also owns any audio it queues for playback. Include provider buffering, interruptions, and ending the call in your application design. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for the Live session lifecycle and [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration) for changes to turn-taking and playback control.

94 

95Keep the provider's call or room identifier alongside the OpenAI session ID so you can trace a conversation across both systems.

96 

97## Next steps with GPT-Live

98 

99- [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=live): connect a server audio stream to GPT-Live.

100- [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live): manage a session from your backend.

101- [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation): connect speech to your reasoning and tool backend.

102- [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations): handle transcripts, session state, and close.

103 

104 

105 

106 

107 

108 

5[SIP](https://en.wikipedia.org/wiki/Session_Initiation_Protocol) is a109[SIP](https://en.wikipedia.org/wiki/Session_Initiation_Protocol) is a

6protocol used to make phone calls over the internet. With SIP and the110protocol used to make phone calls over the internet. With SIP and the

7Realtime API you can direct incoming phone calls to the API.111Realtime API you can direct incoming phone calls to the API.


126The WebSocket behaves exactly like any other Realtime API connection. Send230The WebSocket behaves exactly like any other Realtime API connection. Send

127[`response.create`](https://developers.openai.com/api/reference/resources/realtime/client-events#response.create),231[`response.create`](https://developers.openai.com/api/reference/resources/realtime/client-events#response.create),

128and other client events to control the call, and listen for server events to232and other client events to control the call, and listen for server events to

129track progress. See [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls)233track progress. See [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=realtime)

130for more information.234for more information.

131 235 

132```javascript236```javascript


355 459 

356Now that you've connected over SIP, use the left navigation or click into these pages to start building your realtime application.460Now that you've connected over SIP, use the left navigation or click into these pages to start building your realtime application.

357 461 

358- [Realtime prompting guide](https://developers.openai.com/api/docs/guides/realtime-models-prompting)462- [Realtime prompting guide](https://developers.openai.com/api/docs/guides/voice-prompting)

359- [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations)463- [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations)

360- [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls)464- [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=realtime)

361- [Managing costs](https://developers.openai.com/api/docs/guides/realtime-costs)465- [Managing costs](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=realtime)

362- [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)466- [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)

363 467 

364### Additional Resources468### Additional Resources

guides/voice-webrtc.md +911 −0 created

Details

1# WebRTC

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.

6 

7 

8 

9## Connect a browser to GPT-Live

10 

11Use WebRTC for browser voice applications. Microphone input and generated speech travel on negotiated media tracks. A data channel carries JSON events for transcripts, session updates, and delegated work.

12 

13Your browser creates a Session Description Protocol (SDP) offer. Your application server exchanges it for an answer with `POST /v1/live/sessions`, using the project API key. Keep the key and session configuration on your trusted server.

14 

15### Before you start

16 

17You need:

18 

19- A project API key with access to GPT-Live.

20- A server runtime for your chosen SDK example. The Node.js example requires Node.js 22.6 or later.

21- A browser with microphone permission, running on HTTPS or localhost.

22 

23The example uses Responses delegation with `gpt-5.6-terra` and hosted web search. For backend instructions and application tools, see [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation). For voice and backend usage, see [Cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live).

24 

25### Understand the connection sequence

26 

271. Request microphone access from a user action and add its tracks to a peer connection.

282. Create the data channel and register event listeners before creating the SDP offer.

293. Set the local description, wait for ICE candidate gathering, and send the offer to your server.

304. Have your server post JSON containing `session` and `transport: { type: "webrtc", sdp: ... }` to OpenAI.

315. Apply the returned SDP answer as the remote description. Wait for `session.started` on the data channel before sending application commands.

32 

33The HTTP request starts the session. **Do not send `session.start` on the data channel.** The `oai-events` string in the example is the data-channel label.

34 

35Creating a WebRTC session with `POST /v1/live/sessions` bills 15 seconds of voice duration during initialization. That amount is credited against duration charges once the session starts running; it is not an extra 15 seconds added to the running session. See [WebRTC initialization charges](https://developers.openai.com/api/docs/guides/voice-latency-cost?api=live#webrtc-initialization-charges) for cost accounting.

36 

37### Create the application server

38 

39Save the server example in a new directory and set `OPENAI_API_KEY` in its environment. For Node.js, use `server.mjs` and install `openai` and `express` with `npm install openai express`. Install the corresponding OpenAI SDK for the other language variants; the Ruby example also uses `webrick`. This example binds to `127.0.0.1`, accepts session requests from `http://localhost:3000`, and serves `index.html` from the directory where you run it.

40 

41Choose a server language below; each variant serves `index.html` and the same `/api/session` endpoint on port 3000. Use an SDK version with Live support. Run only one variant at a time.

42 

43```javascript

44import express from "express";

45import OpenAI from "openai";

46import { readFile } from "node:fs/promises";

47import { resolve } from "node:path";

48 

49const app = express();

50const client = new OpenAI({ maxRetries: 0 });

51const port = 3000;

52const origin = `http://localhost:${port}`;

53const indexPath = resolve("index.html");

54 

55app.use(express.json({ limit: "64kb" }));

56app.get("/", async (_request, response) => {

57 response.type("html").send(await readFile(indexPath, "utf8"));

58});

59 

60// Local-only demo. Add your application's authentication and authorization

61// before exposing session creation to other users.

62app.post("/api/session", async (request, response) => {

63 if (request.headers.origin !== origin) {

64 response.status(403).json({ error: "Unexpected request origin" });

65 return;

66 }

67 if (typeof request.body?.sdp !== "string" || !request.body.sdp.trim()) {

68 response.status(400).json({ error: "An SDP offer is required" });

69 return;

70 }

71 if (!process.env.OPENAI_API_KEY) {

72 response.status(503).json({ error: "Set OPENAI_API_KEY on the server" });

73 return;

74 }

75 

76 try {

77 const result = await client.live.create({

78 session: {

79 model: "gpt-live-1",

80 instructions:

81 "Be concise. Delegate requests needing current information to the backend, which can search the web.",

82 delegation: {

83 type: "responses",

84 responses: {

85 model: "gpt-5.6-terra",

86 instructions:

87 "Use web search when current facts are needed. Return concise, grounded results for a spoken conversation.",

88 tools: [{ type: "web_search" }],

89 tool_choice: "auto",

90 },

91 },

92 },

93 transport: {

94 type: "webrtc",

95 sdp: request.body.sdp,

96 },

97 });

98 // Preserve the SDK's typed session ID and SDP answer.

99 response.status(201).json(result);

100 } catch (error) {

101 if (!(error instanceof OpenAI.APIError)) throw error;

102 console.error("Live session creation failed", error.status);

103 response

104 .status(error.status ?? 502)

105 .json({ error: "Live session creation failed" });

106 }

107});

108 

109app.listen(port, "127.0.0.1", () => console.log(`Open ${origin}`));

110```

111 

112```python

113from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

114from pathlib import Path

115 

116from openai import APIError, OpenAI

117from openai.types.live.media_session_config_param import MediaSessionConfigParam

118from pydantic import BaseModel, Field, ValidationError

119 

120client = OpenAI(max_retries=0)

121origin = "http://localhost:3000"

122 

123 

124class SDPOffer(BaseModel):

125 sdp: str = Field(min_length=1)

126 

127 

128class SessionHandler(BaseHTTPRequestHandler):

129 def reply(

130 self, status: int, body: bytes, content_type: str = "application/json"

131 ) -> None:

132 self.send_response(status)

133 self.send_header("Content-Type", content_type)

134 self.end_headers()

135 self.wfile.write(body)

136 

137 def do_GET(self) -> None:

138 if self.path != "/":

139 self.reply(404, b"{}")

140 return

141 self.reply(200, Path("index.html").read_bytes(), "text/html")

142 

143 def do_POST(self) -> None:

144 # Local-only demo: add application authentication before exposing it.

145 if self.path != "/api/session":

146 self.reply(404, b"{}")

147 return

148 if self.headers.get("Origin") != origin:

149 self.reply(403, b'{"error":"Unexpected request origin"}')

150 return

151 length = int(self.headers.get("Content-Length", "0"))

152 if not 0 < length <= 65536:

153 self.reply(400, b'{"error":"An SDP offer is required"}')

154 return

155 try:

156 offer = SDPOffer.model_validate_json(self.rfile.read(length))

157 if not offer.sdp.strip():

158 self.reply(400, b'{"error":"An SDP offer is required"}')

159 return

160 except ValidationError:

161 self.reply(400, b'{"error":"An SDP offer is required"}')

162 return

163 session: MediaSessionConfigParam = {

164 "model": "gpt-live-1",

165 "instructions": "Be concise. Delegate requests needing current information to the backend, which can search the web.",

166 "delegation": {

167 "type": "responses",

168 "responses": {

169 "model": "gpt-5.6-terra",

170 "instructions": "Use web search when current facts are needed. Return concise, grounded results for a spoken conversation.",

171 "tools": [{"type": "web_search"}],

172 "tool_choice": "auto",

173 },

174 },

175 }

176 try:

177 result = client.live.create(

178 session=session, transport={"type": "webrtc", "sdp": offer.sdp}

179 )

180 except APIError as error:

181 self.reply(

182 getattr(error, "status_code", 502),

183 b'{"error":"Live session creation failed"}',

184 )

185 return

186 # Return the SDK's typed session ID and SDP answer unchanged.

187 self.reply(201, result.model_dump_json().encode())

188 

189 

190if __name__ == "__main__":

191 print(f"Open {origin}")

192 ThreadingHTTPServer(("127.0.0.1", 3000), SessionHandler).serve_forever()

193```

194 

195```go

196package main

197 

198import (

199 "encoding/json"

200 "log"

201 "net/http"

202 "os"

203 "strings"

204 

205 "github.com/openai/openai-go/v3"

206 "github.com/openai/openai-go/v3/live"

207 "github.com/openai/openai-go/v3/option"

208)

209 

210func main() {

211 client := openai.NewClient(option.WithMaxRetries(0))

212 const origin = "http://localhost:3000"

213 mux := http.NewServeMux()

214 mux.HandleFunc("GET /{$}", func(w http.ResponseWriter, r *http.Request) {

215 page, err := os.ReadFile("index.html")

216 if err != nil {

217 http.Error(w, "index.html unavailable", 500)

218 return

219 }

220 w.Header().Set("Content-Type", "text/html")

221 w.Write(page)

222 })

223 mux.HandleFunc("POST /api/session", func(w http.ResponseWriter, r *http.Request) {

224 // Local-only demo: add application authentication before exposing it.

225 if r.Header.Get("Origin") != origin {

226 http.Error(w, "Unexpected request origin", 403)

227 return

228 }

229 var offer struct {

230 SDP string `json:"sdp"`

231 }

232 r.Body = http.MaxBytesReader(w, r.Body, 65536)

233 if err := json.NewDecoder(r.Body).Decode(&offer); err != nil || strings.TrimSpace(offer.SDP) == "" {

234 http.Error(w, "An SDP offer is required", 400)

235 return

236 }

237 result, err := client.Live.New(r.Context(), live.LiveNewParams{

238 Session: live.MediaSessionConfigParam{

239 Model: "gpt-live-1",

240 Instructions: openai.String("Be concise. Delegate requests needing current information to the backend, which can search the web."),

241 Delegation: live.MediaSessionConfigDelegationUnionParam{

242 OfResponses: &live.MediaSessionConfigDelegationResponsesParam{

243 Responses: live.ResponsesDelegationConfigParam{

244 Model: "gpt-5.6-terra",

245 Instructions: openai.String("Use web search when current facts are needed. Return concise, grounded results for a spoken conversation."),

246 Tools: []live.ResponsesDelegationConfigToolUnionParam{{

247 OfWebSearch: &live.ResponsesDelegationConfigToolWebSearchParam{},

248 }},

249 ToolChoice: live.ResponsesDelegationConfigToolChoiceUnionParam{OfLiveToolChoiceEnum: openai.String("auto")},

250 },

251 },

252 },

253 },

254 Transport: live.LiveNewParamsTransport{Sdp: offer.SDP},

255 })

256 if err != nil {

257 log.Print(err)

258 http.Error(w, "Live session creation failed", 502)

259 return

260 }

261 // Return the SDK's typed session ID and SDP answer unchanged.

262 w.Header().Set("Content-Type", "application/json")

263 w.WriteHeader(http.StatusCreated)

264 json.NewEncoder(w).Encode(result)

265 })

266 log.Printf("Open %s", origin)

267 log.Fatal(http.ListenAndServe("127.0.0.1:3000", mux))

268}

269```

270 

271```java

272import com.fasterxml.jackson.databind.json.JsonMapper;

273import com.openai.client.OpenAIClient;

274import com.openai.client.okhttp.OpenAIOkHttpClient;

275import com.openai.core.ObjectMappers;

276import com.openai.models.live.LiveCreateParams;

277import com.openai.models.live.LiveCreateResponse;

278import com.openai.models.live.MediaSessionConfig;

279import com.openai.models.live.ResponsesDelegationConfig;

280import com.sun.net.httpserver.HttpExchange;

281import com.sun.net.httpserver.HttpServer;

282import java.io.IOException;

283import java.net.InetSocketAddress;

284import java.nio.charset.StandardCharsets;

285import java.nio.file.Files;

286import java.nio.file.Path;

287 

288public class LiveConnectionWebrtcExample {

289 record SDPOffer(String sdp) {}

290 

291 static void reply(HttpExchange exchange, int status, byte[] body, String contentType)

292 throws IOException {

293 exchange.getResponseHeaders().set("Content-Type", contentType);

294 exchange.sendResponseHeaders(status, body.length);

295 try (var output = exchange.getResponseBody()) {

296 output.write(body);

297 }

298 }

299 

300 public static void main(String[] args) throws IOException {

301 OpenAIClient client = OpenAIOkHttpClient.builder().fromEnv().maxRetries(0).build();

302 JsonMapper json = ObjectMappers.jsonMapper();

303 String origin = "http://localhost:3000";

304 HttpServer server = HttpServer.create(new InetSocketAddress("127.0.0.1", 3000), 0);

305 server.createContext(

306 "/",

307 exchange -> {

308 String path = exchange.getRequestURI().getPath();

309 if (path.equals("/") && exchange.getRequestMethod().equals("GET")) {

310 reply(exchange, 200, Files.readAllBytes(Path.of("index.html")), "text/html");

311 return;

312 }

313 if (!path.equals("/api/session") || !exchange.getRequestMethod().equals("POST")) {

314 reply(exchange, 404, new byte[0], "text/plain");

315 return;

316 }

317 // Local-only demo: add application authentication before exposing it.

318 if (!origin.equals(exchange.getRequestHeaders().getFirst("Origin"))) {

319 reply(exchange, 403, new byte[0], "text/plain");

320 return;

321 }

322 SDPOffer offer;

323 try {

324 byte[] body = exchange.getRequestBody().readNBytes(65537);

325 if (body.length > 65536) throw new IOException("SDP offer too large");

326 offer = json.readValue(body, SDPOffer.class);

327 if (offer.sdp() == null || offer.sdp().isBlank())

328 throw new IOException("Missing SDP offer");

329 } catch (IOException error) {

330 reply(

331 exchange,

332 400,

333 "An SDP offer is required".getBytes(StandardCharsets.UTF_8),

334 "text/plain");

335 return;

336 }

337 try {

338 LiveCreateResponse result =

339 client

340 .live()

341 .create(

342 LiveCreateParams.builder()

343 .session(

344 MediaSessionConfig.builder()

345 .model("gpt-live-1")

346 .instructions(

347 "Be concise. Delegate requests needing current information to the backend, which can search the web.")

348 .responsesDelegation(

349 ResponsesDelegationConfig.builder()

350 .model("gpt-5.6-terra")

351 .instructions(

352 "Use web search when current facts are needed. Return concise, grounded results for a spoken conversation.")

353 .addToolWebSearch()

354 .toolChoice(

355 ResponsesDelegationConfig.ToolChoice

356 .LiveToolChoiceEnum.AUTO)

357 .build())

358 .build())

359 .transport(

360 LiveCreateParams.Transport.builder().sdp(offer.sdp()).build())

361 .build());

362 // Return the SDK's typed session ID and SDP answer unchanged.

363 reply(exchange, 201, json.writeValueAsBytes(result), "application/json");

364 } catch (com.openai.errors.OpenAIException error) {

365 System.err.println(error.getMessage());

366 reply(

367 exchange,

368 502,

369 "Live session creation failed".getBytes(StandardCharsets.UTF_8),

370 "text/plain");

371 }

372 });

373 System.out.println("Open " + origin);

374 server.start();

375 }

376}

377```

378 

379```ruby

380require "json"

381require "openai"

382require "webrick"

383 

384client = OpenAI::Client.new(max_retries: 0)

385origin = "http://localhost:3000"

386server = WEBrick::HTTPServer.new(Port: 3000, BindAddress: "127.0.0.1")

387server.mount_proc("/") do |request, response|

388 if request.path == "/" && request.request_method == "GET"

389 response["Content-Type"] = "text/html"

390 response.body = File.read("index.html")

391 next

392 end

393 unless request.path == "/api/session" && request.request_method == "POST"

394 response.status = 404

395 next

396 end

397 # Local-only demo: add application authentication before exposing it.

398 unless request["Origin"] == origin

399 response.status = 403

400 next

401 end

402 begin

403 raise ArgumentError if request.body.to_s.bytesize > 65_536

404 

405 offer = JSON.parse(request.body.to_s)

406 sdp = offer.fetch("sdp")

407 raise ArgumentError unless sdp.is_a?(String) && !sdp.strip.empty?

408 rescue JSON::ParserError, KeyError, ArgumentError

409 response.status = 400

410 response.body = "An SDP offer is required"

411 next

412 end

413 session = OpenAI::Models::Live::MediaSessionConfig.new(

414 model: "gpt-live-1",

415 instructions: "Be concise. Delegate requests needing current information to the backend, which can search the web.",

416 delegation: OpenAI::Models::Live::MediaSessionConfig::Delegation::Responses.new(

417 responses: OpenAI::Models::Live::ResponsesDelegationConfig.new(

418 model: "gpt-5.6-terra",

419 instructions: "Use web search when current facts are needed. Return concise, grounded results for a spoken conversation.",

420 tools: [OpenAI::Models::Live::ResponsesDelegationConfig::Tool::WebSearch.new],

421 tool_choice: :auto

422 )

423 )

424 )

425 begin

426 result = client.live.create(

427 session: session,

428 transport: OpenAI::Models::Live::LiveCreateParams::Transport.new(sdp: sdp)

429 )

430 # Return the SDK's typed session ID and SDP answer unchanged.

431 response.status = 201

432 response["Content-Type"] = "application/json"

433 response.body = result.to_json

434 rescue OpenAI::Errors::APIError => error

435 warn(error.message)

436 response.status = 502

437 response.body = "Live session creation failed"

438 end

439end

440trap("INT") { server.shutdown }

441puts "Open #{origin}"

442server.start

443```

444 

445 

446Before making the server accessible to other users, protect `/api/session` with your application's authentication, authorization, request limits, and HTTPS. The origin check in this local example does not authenticate users.

447 

448### Create the browser client

449 

450Create `index.html` in the directory where you run the server:

451 

452```html

453<!doctype html>

454<html lang="en">

455 <head>

456 <meta charset="utf-8" />

457 <meta name="viewport" content="width=device-width, initial-scale=1" />

458 <title>GPT-Live connection</title>

459 </head>

460 <body>

461 <script type="module">

462 // Paste the browser code below here.

463 </script>

464 </body>

465</html>

466```

467 

468Paste the following code inside the module script. It adds start and end controls, connects the microphone and output audio, and handles session events. `/api/session` is a route on your application server.

469 

470```javascript

471const start = document.createElement("button");

472start.textContent = "Start conversation";

473const stop = document.createElement("button");

474stop.textContent = "End conversation";

475stop.disabled = true;

476const status = document.createElement("p");

477const audio = new Audio();

478audio.autoplay = true;

479audio.controls = true;

480document.body.append(start, stop, status, audio);

481 

482/** @type {RTCPeerConnection | undefined} */

483let peer;

484/** @type {RTCDataChannel | undefined} */

485let events;

486/** @type {MediaStream | undefined} */

487let microphone;

488/** @type {ReturnType<typeof setTimeout> | undefined} */

489let closeTimeout;

490let ready = false;

491let finalized = false;

492 

493function cleanup() {

494 clearTimeout(closeTimeout);

495 microphone?.getTracks().forEach((track) => track.stop());

496 events?.close();

497 peer?.close();

498 audio.srcObject = null;

499 ready = false;

500 start.disabled = false;

501 stop.disabled = true;

502}

503 

504start.addEventListener("click", async () => {

505 start.disabled = true;

506 finalized = false;

507 status.textContent = "Connecting…";

508 try {

509 const connection = new RTCPeerConnection();

510 peer = connection;

511 connection.addEventListener("track", (event) => {

512 audio.srcObject = new MediaStream([event.track]);

513 audio.play().catch(() => {

514 status.textContent =

515 "Select play on the audio controls to hear the assistant.";

516 });

517 });

518 microphone = await navigator.mediaDevices.getUserMedia({ audio: true });

519 for (const track of microphone.getAudioTracks()) {

520 connection.addTrack(track, microphone);

521 }

522 

523 // Create the event channel before creating the SDP offer.

524 events = connection.createDataChannel("oai-events");

525 events.addEventListener("message", ({ data }) => {

526 /** @type {import("openai/resources/live/live").ServerEvent} */

527 const event = JSON.parse(data);

528 if (event.type === "session.started") {

529 ready = true;

530 stop.disabled = false;

531 status.textContent = "Connected: " + event.session.id;

532 } else if (event.type === "session.closed") {

533 finalized = true;

534 console.log("Final session usage", event.usage);

535 status.textContent = "Conversation ended.";

536 cleanup();

537 } else {

538 // Save transcript and nested Responses events as needed by your app.

539 console.log(event);

540 }

541 });

542 events.addEventListener("close", (event) => {

543 if (event.target !== events) return;

544 if (!finalized) {

545 status.textContent = "Disconnected without final session usage.";

546 cleanup();

547 }

548 });

549 

550 const offer = await connection.createOffer();

551 await connection.setLocalDescription(offer);

552 if (connection.iceGatheringState !== "complete") {

553 await new Promise((resolve, reject) => {

554 const timeout = setTimeout(() => {

555 connection.removeEventListener("icegatheringstatechange", onState);

556 reject(new Error("Timed out while gathering ICE candidates"));

557 }, 10_000);

558 function onState() {

559 if (connection.iceGatheringState !== "complete") return;

560 clearTimeout(timeout);

561 connection.removeEventListener("icegatheringstatechange", onState);

562 resolve(undefined);

563 }

564 connection.addEventListener("icegatheringstatechange", onState);

565 onState();

566 });

567 }

568 

569 const sdp = connection.localDescription?.sdp;

570 if (!sdp) throw new Error("Missing local SDP offer");

571 const response = await fetch("/api/session", {

572 method: "POST",

573 headers: { "Content-Type": "application/json" },

574 body: JSON.stringify({ sdp }),

575 });

576 if (!response.ok) throw new Error(await response.text());

577 /** @type {import("openai/resources/live/live").LiveCreateResponse} */

578 const result = await response.json();

579 console.log("Created session", result.session.id);

580 await connection.setRemoteDescription({

581 type: "answer",

582 sdp: result.transport.sdp,

583 });

584 // The HTTP request started this session. Do not send session.start here.

585 } catch (error) {

586 status.textContent =

587 error instanceof Error ? error.message : String(error);

588 cleanup();

589 }

590});

591 

592stop.addEventListener("click", () => {

593 if (!ready || !events || events.readyState !== "open") return;

594 stop.disabled = true;

595 status.textContent = "Finishing the conversation…";

596 // The session.closed handler is already registered. Keep media and events

597 // alive while pending work drains; only clean up after the final event.

598 events.send(JSON.stringify({ type: "session.close" }));

599 closeTimeout = setTimeout(() => {

600 status.textContent = "Incomplete finalization: no session.closed event.";

601 cleanup();

602 }, 15_000);

603});

604```

605 

606 

607Run your chosen server (`node server.mjs`, `python server.py`, `go run main.go`, `ruby server.rb`, or the Java `LiveConnectionWebrtcExample` class), open `http://localhost:3000`, and select **Start conversation**. After the status changes to **Connected**, ask a question that needs current information to exercise hosted search. Use the audio controls if your browser blocks autoplay.

608 

609### Read the session response

610 

611A successful request returns HTTP 201 with JSON containing the session ID and SDP answer:

612 

613```json

614{

615 "session": { "id": "live_123" },

616 "transport": { "type": "webrtc", "sdp": "<SDP answer>" }

617}

618```

619 

620Read `result.session.id` and pass `result.transport.sdp` to `setRemoteDescription`. Treat the session ID as opaque and preserve it unchanged, including its prefix.

621 

622### Handle media and events

623 

624Send microphone audio and receive generated speech through the media tracks. WebRTC negotiates the audio format through SDP, so omit `audio.format` from session configuration. Do not send `session.input_audio.append` or expect `session.output_audio.delta` on the data channel.

625 

626Use the data channel for transcript deltas, session commands, and nested `response.event` messages. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for transcript handling and lifecycle events, and [Server-side controls](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) if your server needs its own event connection.

627 

628To end the conversation, send `session.close` and keep receiving until `session.closed` before closing the peer connection and microphone tracks. The example registers the final-event listener before sending the command. If the connection fails or times out first, final usage is unconfirmed. See [Usage and graceful close](https://developers.openai.com/api/docs/guides/live-conversations#usage-and-graceful-close) for final usage handling.

629 

630

631 

632

633 

634 

635[WebRTC](https://webrtc.org/) is a powerful set of standard interfaces for building real-time applications. The OpenAI Realtime API supports connecting to realtime models through a WebRTC peer connection.

636 

637For browser-based speech-to-speech voice applications, we recommend starting with [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents), which covers the Agents SDK's higher-level helpers and APIs for managing Realtime sessions. The WebRTC interface is powerful and flexible, but lower level than the Agents SDK.

638 

639When connecting to a Realtime model from the client (like a web browser or

640 mobile device), we recommend using WebRTC rather than WebSockets for more

641 consistent performance.

642 

643For more guidance on building user interfaces on top of WebRTC, [refer to the docs on MDN](https://developer.mozilla.org/en-US/docs/Web/API/WebRTC_API).

644 

645## Overview

646 

647The Realtime API supports two mechanisms for connecting to the Realtime API from the browser, either using ephemeral API keys ([generated via the OpenAI REST API](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets)), or via the new unified interface. Generally, using the unified interface is simpler, but puts your application server in the critical path for session initialization.

648 

649### Connecting using the unified interface

650 

651The process for initializing a WebRTC connection using the unified interface is as follows (assuming a web browser client):

652 

6531. The browser makes a request to a developer-controlled server using the SDP data from its WebRTC peer connection.

6542. The server combines that SDP with its session configuration in a multipart form and sends that to the OpenAI Realtime API, authenticating it with its [standard API key](https://platform.openai.com/settings/organization/api-keys).

655 

656#### Creating a session via the unified interface

657 

658To create a realtime API session via the unified interface, you will need to build a small server-side application (or integrate with an existing one) to make a request to `/v1/realtime/calls`. You will use a [standard API key](https://platform.openai.com/settings/organization/api-keys) to authenticate this request on your backend server.

659 

660Below is an example of a simple Node.js [express](https://expressjs.com/) server which creates a realtime API session:

661 

662```javascript

663import express from "express";

664 

665const app = express();

666 

667// Parse raw SDP payloads posted from the browser

668app.use(express.text({ type: ["application/sdp", "text/plain"] }));

669 

670const sessionConfig = JSON.stringify({

671 type: "realtime",

672 model: "gpt-realtime-2.1",

673 audio: { output: { voice: "marin" } },

674});

675 

676// An endpoint which creates a Realtime API session.

677app.post("/session", async (req, res) => {

678 const fd = new FormData();

679 fd.set("sdp", req.body);

680 fd.set("session", sessionConfig);

681 

682 try {

683 const r = await fetch("https://api.openai.com/v1/realtime/calls", {

684 method: "POST",

685 headers: {

686 Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,

687 "OpenAI-Safety-Identifier": "hashed-user-id",

688 },

689 body: fd,

690 });

691 // Send back the SDP we received from the OpenAI REST API

692 const sdp = await r.text();

693 res.send(sdp);

694 } catch (error) {

695 console.error("Token generation error:", error);

696 res.status(500).json({ error: "Failed to generate token" });

697 }

698});

699 

700app.listen(3000);

701```

702 

703 

704If your application assigns a [safety identifier](https://developers.openai.com/api/docs/guides/safety-best-practices#implement-safety-identifiers)

705for each end user, include it as the `OpenAI-Safety-Identifier` header in this

706server-side request. Use a stable, privacy-preserving value, such as a hashed

707internal user ID. The header should be set by your trusted backend, not by the

708browser.

709 

710#### Connecting to the server

711 

712In the browser, you can use standard WebRTC APIs to connect to the Realtime API via your application server. The client directly POSTs its SDP data to your server.

713 

714```javascript

715// Create a peer connection

716const pc = new RTCPeerConnection();

717 

718// Set up to play remote audio from the model

719audioElement.current = document.createElement("audio");

720audioElement.current.autoplay = true;

721pc.ontrack = (e) => (audioElement.current.srcObject = e.streams[0]);

722 

723// Add local audio track for microphone input in the browser

724const ms = await navigator.mediaDevices.getUserMedia({

725 audio: true,

726});

727pc.addTrack(ms.getTracks()[0]);

728 

729// Set up data channel for sending and receiving events

730const dc = pc.createDataChannel("oai-events");

731 

732// Start the session using the Session Description Protocol (SDP)

733const offer = await pc.createOffer();

734await pc.setLocalDescription(offer);

735 

736const sdpResponse = await fetch("/session", {

737 method: "POST",

738 body: offer.sdp,

739 headers: {

740 "Content-Type": "application/sdp",

741 },

742});

743 

744const answer = {

745 type: "answer",

746 sdp: await sdpResponse.text(),

747};

748await pc.setRemoteDescription(answer);

749```

750 

751 

752### Connecting using an ephemeral token

753 

754The process for initializing a WebRTC connection using an ephemeral API key is as follows (assuming a web browser client):

755 

7561. The browser makes a request to a developer-controlled server to mint an ephemeral API key.

7571. The developer's server uses a [standard API key](https://platform.openai.com/settings/organization/api-keys) to request an ephemeral key from the [OpenAI REST API](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets), and returns that new key to the browser.

7581. The browser uses the ephemeral key to authenticate a session directly with the OpenAI Realtime API as a [WebRTC peer connection](https://developer.mozilla.org/en-US/docs/Web/API/RTCPeerConnection).

759 

760![connect to realtime via WebRTC](https://openaidevs.retool.com/api/file/55b47800-9aaf-48b9-90d5-793ab227ddd3)

761 

762#### Creating an ephemeral token

763 

764To create an ephemeral token to use on the client-side, you will need to build a small server-side application (or integrate with an existing one) to make an [OpenAI REST API](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets) request for an ephemeral key. You will use a [standard API key](https://platform.openai.com/settings/organization/api-keys) to authenticate this request on your backend server.

765 

766Below is an example of a simple Node.js [express](https://expressjs.com/) server which mints an ephemeral API key using the REST API:

767 

768```javascript

769import express from "express";

770 

771const app = express();

772 

773const sessionConfig = JSON.stringify({

774 session: {

775 type: "realtime",

776 model: "gpt-realtime-2.1",

777 audio: {

778 output: {

779 voice: "marin",

780 },

781 },

782 },

783});

784 

785// An endpoint which would work with the client code above - it returns

786// the contents of a REST API request to this protected endpoint

787app.get("/token", async (req, res) => {

788 try {

789 const response = await fetch(

790 "https://api.openai.com/v1/realtime/client_secrets",

791 {

792 method: "POST",

793 headers: {

794 Authorization: `Bearer ${apiKey}`,

795 "Content-Type": "application/json",

796 "OpenAI-Safety-Identifier": "hashed-user-id",

797 },

798 body: sessionConfig,

799 }

800 );

801 

802 const data = await response.json();

803 res.json(data);

804 } catch (error) {

805 console.error("Token generation error:", error);

806 res.status(500).json({ error: "Failed to generate token" });

807 }

808});

809 

810app.listen(3000);

811```

812 

813 

814You can create a server endpoint like this one on any platform that can send and receive HTTP requests. Just ensure that **you only use standard OpenAI API keys on the server, not in the browser.**

815 

816When using ephemeral tokens, set `OpenAI-Safety-Identifier` on the server-side

817request that creates the client secret. The Realtime API binds the identifier to

818the resulting ephemeral token, so the browser does not need to send the safety

819identifier when it later connects with that token.

820 

821#### Connecting to the server

822 

823In the browser, you can use standard WebRTC APIs to connect to the Realtime API with an ephemeral token. The client first fetches a token from your server endpoint, and then POSTs its SDP data (with the ephemeral token) to the Realtime API.

824 

825```javascript

826// Get a session token for OpenAI Realtime API

827const tokenResponse = await fetch("/token");

828const data = await tokenResponse.json();

829const EPHEMERAL_KEY = data.value;

830 

831// Create a peer connection

832const pc = new RTCPeerConnection();

833 

834// Set up to play remote audio from the model

835audioElement.current = document.createElement("audio");

836audioElement.current.autoplay = true;

837pc.ontrack = (e) => (audioElement.current.srcObject = e.streams[0]);

838 

839// Add local audio track for microphone input in the browser

840const ms = await navigator.mediaDevices.getUserMedia({

841 audio: true,

842});

843pc.addTrack(ms.getTracks()[0]);

844 

845// Set up data channel for sending and receiving events

846const dc = pc.createDataChannel("oai-events");

847 

848// Start the session using the Session Description Protocol (SDP)

849const offer = await pc.createOffer();

850await pc.setLocalDescription(offer);

851 

852const sdpResponse = await fetch("https://api.openai.com/v1/realtime/calls", {

853 method: "POST",

854 body: offer.sdp,

855 headers: {

856 Authorization: `Bearer ${EPHEMERAL_KEY}`,

857 "Content-Type": "application/sdp",

858 },

859});

860 

861const answer = {

862 type: "answer",

863 sdp: await sdpResponse.text(),

864};

865await pc.setRemoteDescription(answer);

866```

867 

868 

869## Sending and receiving events

870 

871Realtime API sessions are managed using a combination of [client-sent events](https://developers.openai.com/api/reference/resources/realtime/client-events#session.update) emitted by you as the developer, and [server-sent events](https://developers.openai.com/api/reference/resources/realtime/server-events#error) created by the Realtime API to indicate session lifecycle events.

872 

873When connecting to a Realtime model via WebRTC, you don't have to handle audio events from the model in the same granular way you must with [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime). The WebRTC peer connection object, if configured as above, will do all that work for you.

874 

875To send and receive other client and server events, you can use the WebRTC peer connection's [data channel](https://developer.mozilla.org/en-US/docs/Web/API/WebRTC_API/Using_data_channels).

876 

877```javascript

878// This is the data channel set up in the browser code above...

879const dc = pc.createDataChannel("oai-events");

880 

881// Listen for server events

882dc.addEventListener("message", (e) => {

883 const event = JSON.parse(e.data);

884 console.log(event);

885});

886 

887// Send client events

888const event = {

889 type: "conversation.item.create",

890 item: {

891 type: "message",

892 role: "user",

893 content: [

894 {

895 type: "input_text",

896 text: "hello there!",

897 },

898 ],

899 },

900};

901dc.send(JSON.stringify(event));

902```

903 

904 

905To learn more about managing Realtime conversations, refer to the [Realtime conversations guide](https://developers.openai.com/api/docs/guides/realtime-conversations).

906 

907[Realtime Console

908 

909 

910 

911 Check out the WebRTC Realtime API in this light weight example app.](https://github.com/openai/openai-realtime-console/)

guides/voice-websockets.md +494 −0 created

Details

1# WebSockets

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5Choose the API your application uses. Each API has its own authentication, session creation, and event contract.

6 

7 

8 

9## Connect a server to GPT-Live

10 

11Use a primary WebSocket when your server captures audio or relays an audio stream for a client. It carries audio and JSON events in both directions. Keep the project API key on that trusted server. For browser and mobile applications, start with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live).

12 

13This guide covers the primary audio connection. A [sideband connection](https://developers.openai.com/api/docs/guides/voice-server-controls?api=live) lets a server observe and control an existing Live session. A [Responses WebSocket](https://developers.openai.com/api/docs/guides/websocket-mode) connects your backend to the Responses API for reasoning and tools. Neither replaces the primary audio connection.

14 

15### Authenticate and start the session

16 

171. Connect to `wss://api.openai.com/v1/live/sessions` with no query parameters. Authenticate with `Authorization: Bearer $OPENAI_API_KEY` and include the connection headers shown in the example.

182. Send `session.start` as the first message. Put the model, conversation instructions, audio format, voice, and delegation configuration inside the `session` object.

193. Wait for `session.started` before sending audio or application commands. It contains the resolved session configuration and session ID.

20 

21The example below uses Marin, PCM16 audio at 24 kHz, and a Responses backend with web search. Keep conversation instructions short. Configure backend instructions, tools, and tool permissions through [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation).

22 

23### Stream audio with an SDK

24 

25For Node.js, install `openai` and `ws` with `npm install openai ws` and save the JavaScript example as `client.mjs`. For Python on macOS or Linux, install `openai[realtime]` and save the Python example as `client.py`. Set `OPENAI_API_KEY` in the server environment. These examples require an SDK version with Live support. The example reads raw, mono PCM16 audio at 24 kHz from standard input and writes returned audio in the same format to standard output. Connect these streams to your application's audio capture and playback. Logs and transcript events go to standard error so they do not corrupt the audio stream.

26 

27```javascript

28import OpenAI from "openai";

29import { LiveWS } from "openai/resources/live/ws";

30 

31// stdin and stdout carry raw mono PCM16 audio at 24 kHz, not WAV files.

32// Supply microphone bytes continuously and play stdout in the same format.

33process.stdin.pause();

34const ws = new LiveWS(new OpenAI());

35let started = false;

36let closing = false;

37let finalized = false;

38let pendingByte = Buffer.alloc(0);

39/** @type {ReturnType<typeof setTimeout> | undefined} */

40let closeTimeout;

41 

42ws.socket.on("open", () => {

43 ws.send({

44 type: "session.start",

45 event_id: "event_start",

46 session: {

47 model: "gpt-live-1",

48 instructions:

49 "Be concise. Delegate requests needing current information to the backend, which can search the web.",

50 audio: {

51 format: { type: "audio/pcm", rate: 24000 },

52 output: { voice: "marin" },

53 },

54 delegation: {

55 type: "responses",

56 responses: {

57 model: "gpt-5.6-luna",

58 tools: [{ type: "web_search" }],

59 tool_choice: "auto",

60 },

61 },

62 },

63 });

64});

65 

66process.stdin.on("data", (chunk) => {

67 if (!started || closing || ws.socket.readyState !== 1) return;

68 const bytes = Buffer.concat([pendingByte, chunk]);

69 const completeLength = bytes.length - (bytes.length % 2);

70 pendingByte = bytes.subarray(completeLength);

71 if (completeLength) {

72 ws.send({

73 type: "session.input_audio.append",

74 audio: bytes.subarray(0, completeLength).toString("base64"),

75 });

76 }

77});

78 

79// Register the final-event handler before any close command can be sent.

80ws.on("event", (event) => {

81 if (event.type === "session.started") {

82 started = true;

83 console.error("Session ready", event.session.id);

84 process.stdin.resume();

85 } else if (event.type === "session.output_audio.delta") {

86 process.stdout.write(Buffer.from(event.delta, "base64"));

87 } else if (event.type === "session.closed") {

88 finalized = true;

89 clearTimeout(closeTimeout);

90 process.stdin.pause();

91 console.error("Final session usage", event.usage);

92 ws.close();

93 } else {

94 // Includes transcript deltas and nested response.event usage.

95 console.error(JSON.stringify(event));

96 }

97});

98 

99process.on("SIGINT", () => {

100 if (closing) return;

101 if (!started || ws.socket.readyState !== 1) {

102 ws.socket.platformSocket.terminate();

103 return;

104 }

105 closing = true;

106 process.stdin.pause();

107 ws.send({ type: "session.close" });

108 closeTimeout = setTimeout(() => {

109 console.error("Incomplete finalization: session.closed was not received");

110 process.exitCode = 1;

111 ws.socket.platformSocket.terminate();

112 }, 15_000);

113});

114ws.on("error", (error) => {

115 console.error(error.message);

116 process.exitCode = 1;

117});

118ws.socket.on("close", () => {

119 clearTimeout(closeTimeout);

120 process.stdin.pause();

121 if (!finalized) {

122 console.error("Connection closed without final session usage");

123 process.exitCode = 1;

124 }

125});

126```

127 

128```python

129import asyncio

130import base64

131import os

132import signal

133import sys

134 

135from openai import AsyncOpenAI

136from openai.types.live.session_config_param import SessionConfigParam

137 

138 

139async def main() -> None:

140 # stdin/stdout carry raw mono PCM16 at 24 kHz, not WAV files.

141 session: SessionConfigParam = {

142 "model": "gpt-live-1",

143 "instructions": "Be concise. Delegate requests needing current information to the backend, which can search the web.",

144 "audio": {

145 "format": {"type": "audio/pcm", "rate": 24000},

146 "output": {"voice": "marin"},

147 },

148 "delegation": {

149 "type": "responses",

150 "responses": {

151 "model": "gpt-5.6-luna",

152 "tools": [{"type": "web_search"}],

153 "tool_choice": "auto",

154 },

155 },

156 }

157 loop = asyncio.get_running_loop()

158 chunks: asyncio.Queue[bytes] = asyncio.Queue()

159 finalized = asyncio.Event()

160 closing = False

161 pending = b""

162 close_task: asyncio.Task[None] | None = None

163 

164 def read_audio() -> None:

165 chunk = os.read(sys.stdin.fileno(), 4800)

166 if chunk:

167 chunks.put_nowait(chunk)

168 else:

169 # EOF does not end a Live session. Use SIGINT to finalize it.

170 loop.remove_reader(sys.stdin.fileno())

171 

172 async with AsyncOpenAI() as client:

173 async with client.live.connect() as connection:

174 

175 async def send_audio() -> None:

176 nonlocal pending

177 while True:

178 chunk = pending + await chunks.get()

179 complete = len(chunk) - len(chunk) % 2

180 pending = chunk[complete:]

181 if complete and not closing:

182 await connection.session.input_audio.append(

183 audio=base64.b64encode(chunk[:complete]).decode("ascii")

184 )

185 

186 async def close_session() -> None:

187 nonlocal closing

188 closing = True

189 loop.remove_reader(sys.stdin.fileno())

190 await connection.session.close()

191 try:

192 await asyncio.wait_for(finalized.wait(), timeout=15)

193 except TimeoutError:

194 print(

195 "Incomplete finalization: no session.closed event",

196 file=sys.stderr,

197 )

198 await connection.close()

199 

200 def request_close() -> None:

201 nonlocal close_task

202 if close_task is None:

203 close_task = asyncio.create_task(close_session())

204 

205 # Start receiving before a close command can be requested.

206 await connection.session.start(session=session, event_id="event_start")

207 sender = asyncio.create_task(send_audio())

208 try:

209 async for event in connection:

210 if event.type == "session.started":

211 print("Session ready", event.session.id, file=sys.stderr)

212 loop.add_reader(sys.stdin.fileno(), read_audio)

213 loop.add_signal_handler(signal.SIGINT, request_close)

214 elif event.type == "session.output_audio.delta":

215 sys.stdout.buffer.write(base64.b64decode(event.delta))

216 sys.stdout.buffer.flush()

217 elif event.type == "session.closed":

218 print("Final session usage", event.usage, file=sys.stderr)

219 finalized.set()

220 break

221 else:

222 print(event.model_dump_json(), file=sys.stderr)

223 finally:

224 loop.remove_reader(sys.stdin.fileno())

225 loop.remove_signal_handler(signal.SIGINT)

226 sender.cancel()

227 await asyncio.gather(sender, return_exceptions=True)

228 if close_task is not None:

229 await close_task

230 if not finalized.is_set():

231 raise RuntimeError("Connection closed without final session usage")

232 

233 

234if __name__ == "__main__":

235 asyncio.run(main())

236```

237 

238 

239Run `node client.mjs` or `python client.py` with your audio source and player attached. After `Session ready` appears, supply a continuous microphone stream paced at its recorded sample rate. Piping an entire file at once does not simulate a live microphone. EOF on the audio source does not end the conversation. Send `SIGINT` to the process to request a graceful close.

240 

241The example connects the audio streams; your application handles capture, buffering, playback, and resampling when needed. Test these parts with your devices and network before evaluating model behavior.

242 

243### Choose the audio format

244 

245Set `session.audio.format` at startup. One format applies to both input and output and cannot change during the session.

246 

247- `{"type":"audio/pcm","rate":24000}`: mono signed 16-bit little-endian PCM at 24 kHz; the default.

248- `{"type":"audio/pcm","rate":16000}`: mono signed 16-bit little-endian PCM at 16 kHz.

249- `{"type":"audio/pcmu","rate":8000}`: G.711 μ-law at 8 kHz, one byte per sample.

250- `{"type":"audio/pcma","rate":8000}`: G.711 A-law at 8 kHz, one byte per sample.

251 

252Base64-encode raw bytes without a WAV or other container header. PCM chunks must contain complete 16-bit samples, so their byte length must be even. The example carries a trailing byte into the next input chunk. Chunk boundaries are otherwise arbitrary: preserve a continuous, ordered stream.

253 

254Resample audio when its sample rate differs from the configured rate. Changing the format setting does not convert your input bytes. To adapt the example for G.711, forward each chunk's codec bytes without the PCM-specific two-byte alignment logic, and configure the output player for the same codec. A matching G.711 stream can pass through without conversion to PCM. See [Telephony integrations](https://developers.openai.com/api/docs/guides/voice-sip?api=live) for connecting a phone call.

255 

256### Send and receive events

257 

258Send each event as a JSON text message. Audio travels as base64 inside those messages.

259 

260- **Send audio:** send `session.input_audio.append` with raw, base64-encoded bytes in `audio`. Audio appends have no acknowledgment.

261- **Receive audio:** decode `delta` from each `session.output_audio.delta` event and queue the audio for playback in order, using the configured format.

262- **Receive transcripts:** append the text in `delta` from `session.input_transcript.delta` and `session.output_transcript.delta` to the corresponding transcript.

263- **Receive backend events:** when using Responses delegation, process the nested `event` in each `response.event` envelope.

264- **Handle errors:** handle rejected commands and session errors from `error` events. Use `error.client_event_id`, when present, to identify the command.

265 

266Output audio events have no timing fields, and GPT-Live does not emit an output-audio-done event. Track your playback queue to know which received audio has played. Transcript timestamps describe intervals on the session timeline; they do not mark audio playback completion. A backend response completing also does not mean the assistant has finished speaking.

267 

268GPT-Live manages when to listen and speak as audio streams. It does not use Realtime's input-buffer commit and `response.create` voice-turn loop. In Live, `response.create` starts or continues delegated backend work. See [Delegation and tools](https://developers.openai.com/api/docs/guides/live-delegation) for that workflow.

269 

270### Configure an ongoing session

271 

272The Live model, initial conversation instructions, audio format, voice, and delegation mode are fixed at startup. Use `session.update` for supported settings within the existing delegation mode; omitted settings keep their current values. A successful update returns `session.updated` with the resolved session configuration.

273 

274Use `session.instructions.append` to add conversation instructions, and `session.input_audio.mute` or `session.input_audio.unmute` to control incoming audio. Muting input does not cancel backend work or stop generated speech. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for context updates, transcripts, input controls, and usage.

275 

276### Close the session

277 

278Send `session.close` when the conversation ends. Install the `session.closed` listener first, keep receiving until that event arrives, and then release the connection. The example waits up to 15 seconds and reports incomplete finalization if the terminal event never arrives.

279 

280Preserve the final voice usage from `session.closed` and the backend usage events already received. Voice-duration updates are cumulative snapshots; do not add them together. A transport failure or timeout before `session.closed` leaves final usage unconfirmed. See [Managing sessions](https://developers.openai.com/api/docs/guides/live-conversations) for the full lifecycle.

281 

282

283 

284

285 

286 

287[WebSockets](https://developer.mozilla.org/en-US/docs/Web/API/WebSockets_API) are a broadly supported API for realtime data transfer, and a great choice for connecting to the OpenAI Realtime API in server-to-server applications. For browser and mobile clients, we recommend connecting via [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime).

288 

289In a server-to-server integration with Realtime, your backend system will connect via WebSocket directly to the Realtime API. You can use a [standard API key](https://platform.openai.com/settings/organization/api-keys) to authenticate this connection, since the token will only be available on your secure backend server.

290 

291![connect directly to realtime API](https://openaidevs.retool.com/api/file/464d4334-c467-4862-901b-d0c6847f003a)

292 

293## Connect via WebSocket

294 

295Below are several examples of connecting via WebSocket to the Realtime API. In addition to using the WebSocket URL below, you will also need to pass an authentication header using your OpenAI API key. If your application assigns [safety identifiers](https://developers.openai.com/api/docs/guides/safety-best-practices#implement-safety-identifiers), pass the stable, privacy-preserving identifier for the end user in the `OpenAI-Safety-Identifier` header.

296 

297It is possible to use WebSocket in browsers with an ephemeral API token as shown in the [WebRTC connection guide](https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime), but if you are connecting from a client like a browser or mobile app, WebRTC will be a more robust solution in most cases.

298 

299<ContentSwitcher

300 id="connection-example"

301 initialValue="ws"

302 options={[

303 { value: "ws", label: "ws module (Node.js)" },

304 { value: "python", label: "websocket-client (Python)" },

305 { value: "ruby", label: "OpenAI SDK (Ruby)" },

306 { value: "websocket", label: "WebSocket (browsers)" },

307 ]}

308>

309

310 

311

312ws module (Node.js)

313 

314 Connect using the ws module (Node.js)

315 

316```javascript

317import WebSocket from "ws";

318 

319const url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1";

320const ws = new WebSocket(url, {

321 headers: {

322 Authorization: "Bearer " + process.env.OPENAI_API_KEY,

323 "OpenAI-Safety-Identifier": "hashed-user-id",

324 },

325});

326 

327ws.on("open", function open() {

328 console.log("Connected to server.");

329});

330 

331ws.on("message", function incoming(message) {

332 console.log(JSON.parse(message.toString()));

333});

334```

335 

336

337 

338

339 

340

341websocket-client (Python)

342 

343 Connect with websocket-client (Python)

344 

345```python

346# example requires websocket-client library:

347# pip install websocket-client

348 

349import os

350import json

351import websocket

352 

353OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]

354 

355url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1"

356headers = [

357 "Authorization: Bearer " + OPENAI_API_KEY,

358 "OpenAI-Safety-Identifier: hashed-user-id",

359]

360 

361 

362def on_open(ws):

363 print("Connected to server.")

364 

365 

366def on_message(ws, message):

367 data = json.loads(message)

368 print("Received event:", json.dumps(data, indent=2))

369 

370 

371ws = websocket.WebSocketApp(

372 url,

373 header=headers,

374 on_open=on_open,

375 on_message=on_message,

376)

377 

378ws.run_forever()

379```

380 

381

382 

383

384 

385

386OpenAI SDK (Ruby)

387 

388

389 

390 Install the required gems with

391 `gem install openai async-websocket`.

392

393 

394 Connect with the OpenAI SDK (Ruby)

395 

396```ruby

397require "openai"

398 

399client = OpenAI::Client.new(

400 default_headers: {"OpenAI-Safety-Identifier" => "hashed-user-id"}

401)

402 

403client.realtime.connect(model: "gpt-realtime-2.1") do |connection|

404 puts("Connected to the Realtime API: #{connection.url.host}")

405 connection.each { |event| puts("Received event: #{event.type}") }

406end

407```

408 

409

410 

411

412 

413

414WebSocket (browsers)

415 

416 Connect with standard WebSocket (browsers)

417 

418```javascript

419/*

420Note that in client-side environments like web browsers, we recommend

421using WebRTC instead. It is possible, however, to use the standard

422WebSocket interface in browser-like environments like Deno and

423Cloudflare Workers.

424*/

425 

426const ws = new WebSocket(

427 "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1",

428 [

429 "realtime",

430 // Use a short-lived token fetched from your application server.

431 "openai-insecure-api-key." + OPENAI_REALTIME_EPHEMERAL_KEY,

432 // Optional

433 "openai-organization." + OPENAI_ORG_ID,

434 "openai-project." + OPENAI_PROJECT_ID,

435 ]

436);

437 

438ws.addEventListener("open", function open() {

439 console.log("Connected to server.");

440});

441 

442ws.addEventListener("message", function incoming(event) {

443 console.log(event.data);

444});

445```

446 

447 

448 

449## Sending and receiving events

450 

451Realtime API sessions are managed using a combination of [client-sent events](https://developers.openai.com/api/reference/resources/realtime/client-events#session.update) emitted by you as the developer, and [server-sent events](https://developers.openai.com/api/reference/resources/realtime/server-events#error) created by the Realtime API to indicate session lifecycle events.

452 

453Over a WebSocket, you will both send and receive JSON-serialized events as strings of text, as in this Node.js example below (the same principles apply for other WebSocket libraries):

454 

455```javascript

456import WebSocket from "ws";

457 

458const url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1";

459const ws = new WebSocket(url, {

460 headers: {

461 Authorization: "Bearer " + process.env.OPENAI_API_KEY,

462 "OpenAI-Safety-Identifier": "hashed-user-id",

463 },

464});

465 

466ws.on("open", function open() {

467 console.log("Connected to server.");

468 

469 // Send client events over the WebSocket once connected

470 ws.send(

471 JSON.stringify({

472 type: "session.update",

473 session: {

474 type: "realtime",

475 instructions: "Be extra nice today!",

476 },

477 })

478 );

479});

480 

481// Listen for and parse server events

482ws.on("message", function incoming(message) {

483 console.log(JSON.parse(message.toString()));

484});

485```

486 

487 

488The WebSocket interface is perhaps the lowest-level interface available to interact with a Realtime model, where you will be responsible for both sending and processing Base64-encoded audio chunks over the socket connection.

489 

490To learn how to send and receive audio over Websockets, refer to the [Realtime conversations guide](https://developers.openai.com/api/docs/guides/realtime-conversations#handling-audio-with-websockets).

491 

492

493 

494</ContentSwitcher>

Details

79| `/v1/batches` | No | 30 days | Until deleted | No | No |79| `/v1/batches` | No | 30 days | Until deleted | No | No |

80| `/v1/moderations` | No | None | None | Yes | No |80| `/v1/moderations` | No | None | None | Yes | No |

81| `/v1/completions` | No | 30 days | None | Yes | No |81| `/v1/completions` | No | 30 days | None | Yes | No |

82| `/v1/live/sessions` | No | 30 days | None, or 30 days if stored | Yes, with limitations below | No |

82| `/v1/realtime` | No | 30 days | None | Yes | No |83| `/v1/realtime` | No | 30 days | None | Yes | No |

83| `/v1/videos` | No | 30 days | None | No | No |84| `/v1/videos` | No | 30 days | None | No | No |

84 85 


241The complete, unfiltered regional support table follows. Model snapshots for each service are listed in **API Endpoint, tool and model support**. When regional processing supports only a subset of snapshots, that subset is included in the processing-services cell.242The complete, unfiltered regional support table follows. Model snapshots for each service are listed in **API Endpoint, tool and model support**. When regional processing supports only a subset of snapshots, that subset is included in the processing-services cell.

242 243 

243| Region | Domain prefix | Regional storage | Regional processing | MAM or ZDR required | Supported modes | Storage services | Processing services |244| Region | Domain prefix | Regional storage | Regional processing | MAM or ZDR required | Supported modes | Storage services | Processing services |

244| -------------------------- | ------------------- | :--------------: | :-----------------: | :-----------------: | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |245| -------------------------- | ------------------- | :--------------: | :-----------------: | :-----------------: | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

245| United States | `us.api.openai.com` | Yes | Yes | No | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`Code Interpreter tool`<br />`File Search`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` |246| United States | `us.api.openai.com` | Yes | Yes | No | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/live/sessions`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/live/sessions`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`Code Interpreter tool`<br />`File Search`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` |

246| Europe (EEA + Switzerland) | `eu.api.openai.com` | Yes | Yes | Yes\*\* | Text, Audio, Voice, Image\* | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`Code Interpreter tool`<br />`File Search`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` |247| Europe (EEA + Switzerland) | `eu.api.openai.com` | Yes | Yes | Yes\*\* | Text, Audio, Voice, Image\* | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/live/sessions`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/evals`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/live/sessions`<br />`/v1/realtime`<br />`/v1/realtime/transcription_sessions`<br />`/v1/realtime/translations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`Code Interpreter tool`<br />`File Search`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` |

247| Australia\* | `au.api.openai.com` | Yes | No | Yes | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | None |248| Australia\* | `au.api.openai.com` | Yes | No | Yes | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | None |

248| Canada\* | `ca.api.openai.com` | Yes | No | Yes | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | None |249| Canada\* | `ca.api.openai.com` | Yes | No | Yes | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | None |

249| Japan\* | `jp.api.openai.com` | Yes | No | Yes | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | None |250| Japan\* | `jp.api.openai.com` | Yes | No | Yes | Text, Audio, Voice, Image | `/v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech`<br />`/v1/batches`<br />`/v1/chat/completions`<br />`/v1/embeddings`<br />`/v1/files`<br />`/v1/fine_tuning/jobs`<br />`/v1/images/edits`<br />`/v1/images/generations`<br />`/v1/moderations`<br />`/v1/responses`<br />`/v1/responses File Search`<br />`/v1/responses Web Search`<br />`/v1/vector_stores`<br />`Code Interpreter tool`<br />`File Search`<br />`File Uploads`<br />`Remote MCP server tool`<br />`Scale Tier`<br />`Structured Outputs (excluding schema)`<br />`Supported input modalities` | None |


271| `/v1/images/edits` | Images | All listed regions | United States, Europe (EEA + Switzerland) | `gpt-image-2.5-sunburst`, `gpt-image-2.5-sunburst-2026-09-08`, `gpt-image-2.5-flare`, `gpt-image-2.5-flare-2026-09-08`, `gpt-image-2`, `gpt-image-1`, `gpt-image-1.5`, `gpt-image-1-mini` | None | — |272| `/v1/images/edits` | Images | All listed regions | United States, Europe (EEA + Switzerland) | `gpt-image-2.5-sunburst`, `gpt-image-2.5-sunburst-2026-09-08`, `gpt-image-2.5-flare`, `gpt-image-2.5-flare-2026-09-08`, `gpt-image-2`, `gpt-image-1`, `gpt-image-1.5`, `gpt-image-1-mini` | None | — |

272| `/v1/images/generations` | Images | All listed regions | United States, Europe (EEA + Switzerland) | `gpt-image-2.5-sunburst`, `gpt-image-2.5-sunburst-2026-09-08`, `gpt-image-2.5-flare`, `gpt-image-2.5-flare-2026-09-08`, `gpt-image-2`, `gpt-image-1`, `gpt-image-1.5`, `gpt-image-1-mini` | None | — |273| `/v1/images/generations` | Images | All listed regions | United States, Europe (EEA + Switzerland) | `gpt-image-2.5-sunburst`, `gpt-image-2.5-sunburst-2026-09-08`, `gpt-image-2.5-flare`, `gpt-image-2.5-flare-2026-09-08`, `gpt-image-2`, `gpt-image-1`, `gpt-image-1.5`, `gpt-image-1-mini` | None | — |

273| `/v1/moderations` | Moderation | All listed regions | United States, Europe (EEA + Switzerland) | `omni-moderation-latest` | None | — |274| `/v1/moderations` | Moderation | All listed regions | United States, Europe (EEA + Switzerland) | `omni-moderation-latest` | None | — |

275| `/v1/live/sessions` | GPT-Live | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-live-1` | None | — |

274| `/v1/realtime` | Realtime | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-realtime`, `gpt-realtime-1.5`, `gpt-realtime-mini`, `gpt-realtime-2`, `gpt-realtime-2.1`, `gpt-realtime-2.1-mini` | None | — |276| `/v1/realtime` | Realtime | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-realtime`, `gpt-realtime-1.5`, `gpt-realtime-mini`, `gpt-realtime-2`, `gpt-realtime-2.1`, `gpt-realtime-2.1-mini` | None | — |

275| `/v1/realtime/transcription_sessions` | Realtime | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-realtime-whisper`, `gpt-live-transcribe`, `gpt-transcribe` | None | — |277| `/v1/realtime/transcription_sessions` | Realtime | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-realtime-whisper`, `gpt-live-transcribe`, `gpt-transcribe` | None | — |

276| `/v1/realtime/translations` | Realtime | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-realtime-translate` | None | — |278| `/v1/realtime/translations` | Realtime | United States, Europe (EEA + Switzerland) | United States, Europe (EEA + Switzerland) | `gpt-realtime-translate` | None | — |


300- Cannot set background=True in EU region.302- Cannot set background=True in EU region.

301- [Extended prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-retention) in regions that do not support Regional processing may require that OpenAI process and temporarily store Customer Content outside of the Region to deliver the services.303- [Extended prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-retention) in regions that do not support Regional processing may require that OpenAI process and temporarily store Customer Content outside of the Region to deliver the services.

302 304 

305#### /v1/live/sessions

306 

307GPT-Live sessions are eligible for Zero Data Retention. With Zero Data Retention enabled, `store` is treated as `false`, even if a request sets it to `true`.

308 

309Session storage is disabled by default. For projects with session storage enabled, `store: true` retains the completed session recording for 30 days so that it can be downloaded or used to start a forked session. Stored sessions and their index expire after 30 days. Recording downloads and forking require a data policy that permits persistence and are not available with Zero Data Retention.

310 

311Setting `store: false` on a fork prevents storage of the new session; it does not delete the source recording or remove the authorization required to read it. The API does not provide a public stored-session deletion endpoint.

312 

313GPT-Live supports data residency in the United States and Europe. Delegated backend models and tools have their own data controls; check the applicable endpoint and feature entries on this page.

314 

303#### /v1/realtime315#### /v1/realtime

304 316 

305Tracing is not currently EU data residency compliant for `/v1/realtime`.317Tracing is not currently EU data residency compliant for `/v1/realtime`.

libraries.md +1 −1

Details

173<dependency>173<dependency>

174 <groupId>com.openai</groupId>174 <groupId>com.openai</groupId>

175 <artifactId>openai-java</artifactId>175 <artifactId>openai-java</artifactId>

176 <version>4.61.0</version>176 <version>4.62.0</version>

177</dependency>177</dependency>

178```178```

179 179 

models.md +1 −0

Details

87- [GPT-Image-2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model87- [GPT-Image-2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model

88- [GPT-Image-2.5 Flare](/api/docs/models/gpt-image-2.5-flare.md): Fast, high-quality everyday image generation88- [GPT-Image-2.5 Flare](/api/docs/models/gpt-image-2.5-flare.md): Fast, high-quality everyday image generation

89- [GPT-Image-2.5 Sunburst](/api/docs/models/gpt-image-2.5-sunburst.md): Our most capable model for image generation and editing89- [GPT-Image-2.5 Sunburst](/api/docs/models/gpt-image-2.5-sunburst.md): Our most capable model for image generation and editing

90- [GPT-Live 1](/api/docs/models/gpt-live-1.md): Our premier model for natural, expressive voice conversations with smooth interruption handling.

90- [GPT-Live-Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription91- [GPT-Live-Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription

91- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU92- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU

92- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency93- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency

models/all.md +1 −0

Details

87- [GPT-Image-2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model87- [GPT-Image-2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model

88- [GPT-Image-2.5 Flare](/api/docs/models/gpt-image-2.5-flare.md): Fast, high-quality everyday image generation88- [GPT-Image-2.5 Flare](/api/docs/models/gpt-image-2.5-flare.md): Fast, high-quality everyday image generation

89- [GPT-Image-2.5 Sunburst](/api/docs/models/gpt-image-2.5-sunburst.md): Our most capable model for image generation and editing89- [GPT-Image-2.5 Sunburst](/api/docs/models/gpt-image-2.5-sunburst.md): Our most capable model for image generation and editing

90- [GPT-Live 1](/api/docs/models/gpt-live-1.md): Our premier model for natural, expressive voice conversations with smooth interruption handling.

90- [GPT-Live-Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription91- [GPT-Live-Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription

91- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU92- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU

92- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency93- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency

quickstart.md +1 −1

Details

190<dependency>190<dependency>

191 <groupId>com.openai</groupId>191 <groupId>com.openai</groupId>

192 <artifactId>openai-java</artifactId>192 <artifactId>openai-java</artifactId>

193 <version>4.61.0</version>193 <version>4.62.0</version>

194</dependency>194</dependency>

195```195```

196 196