guides/audio.md +29 −497
11# Audio and speech# Audio and voice
2 2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4 4
55Audio models can understand spoken input, generate spoken output, or do both in the same interaction. This guide explains the vocabulary used across OpenAI's audio docs. When you're ready to choose an implementation path, start with the [Realtime and audio overview](https://developers.openai.com/api/docs/guides/realtime).For a new conversational voice application, start with **[GPT-Live](https://developers.openai.com/api/docs/guides/live)**. It can listen while speaking and keep the conversation moving while a backend agent reasons, uses tools, or completes a task.
6 6
77## Audio modalitiesConnect your first conversation with the [WebRTC quickstart](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live), then write a short [Live prompt](https://developers.openai.com/api/docs/guides/live-prompting). If you already have a Realtime application or text agent, follow [Migrate to GPT-Live](https://developers.openai.com/api/docs/guides/live-migration).
8 8
99An audio application combines one or more of these modalities:## Choose another audio workflow
10 10
11| Modality | Meaning | Common use cases |
12| --------------- | -------------------------------------------- | ------------------------------------------------- |
13| Audio input | The model receives sound from a user or app. | Voice agents, transcription, translation. |
14| Audio output | The model or API returns spoken audio. | Voice agents, text to speech, spoken responses. |
15| Text transcript | Speech becomes text. | Captions, call analysis, search, records. |
16| Text prompt | Text controls what the model says or does. | Speech generation, scripted voice flows, prompts. |
17 11
18## Common speech tasks
19 12
20**Speech to text** converts speech into text. Use it for captions, notes, transcripts, analytics, search, and accessibility. Transcription can be request-based for files or streaming for live audio. Start with the [Transcription overview](https://developers.openai.com/api/docs/guides/transcription) to choose a workflow and model.
21 13
22**Text to speech** converts text into spoken audio. Use it for narration, assistants, accessibility, and generated voice responses. Speech generation can stream audio back as the model produces it.
23 14
24**Speech to speech** lets a model listen, reason, and speak in one low-latency session. Use it for conversational voice agents when the assistant needs to respond, call tools, or maintain session state.
25 15
26**Speech translation** listens to speech in one language and returns translated speech or transcript output in another language. Use a dedicated realtime translation session when translation should begin continuously as audio arrives.
27 16
28## Streaming and latency
29 17
30Streaming means the client and service exchange partial input or output while the interaction is still active. Streaming is useful when users expect immediate feedback, such as live captions, calls, voice agents, and translation.
31 18
32Lower latency requires a realtime connection, more careful audio handling, and a session model that can emit partial events. Request-based APIs are simpler for file uploads and non-interactive work, but they don't support the same live interaction patterns.
33 19
34## Request-based APIs and realtime sessions
35 20
36OpenAI supports two broad audio architectures:
37 21
38| Architecture | Use when | Examples |
39| --------------------------- | ---------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
40| Request-based audio APIs | You have a file, a text input, or a bounded request. | [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text), [text to speech](https://developers.openai.com/api/docs/guides/text-to-speech). |
41| Realtime sessions | Audio is live and the app needs low-latency events. | [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents), [translation](https://developers.openai.com/api/docs/guides/realtime-translation), [transcription](https://developers.openai.com/api/docs/guides/realtime-transcription). |
42| Multimodal Chat Completions | You are extending an existing chat flow with audio. | [Audio input or output](#add-audio-to-your-existing-application). |
43 22
44For build-path guidance, see the [Realtime and audio overview](https://developers.openai.com/api/docs/guides/realtime).
45 23
46## Add audio to your existing application
47 24
48Models such as [`gpt-realtime-2.1`](https://developers.openai.com/api/docs/models/gpt-realtime-2.1) and [`gpt-audio-1.5`](https://developers.openai.com/api/docs/models/gpt-audio-1.5) are natively multimodal, meaning they can understand and generate audio and text as input and output.
49 25
50For live browser speech-to-speech interactions, start with a realtime session in the Agents SDK for JavaScript:
51 26
5227Start a realtime voice sessionUse the Realtime API when you need its session and tool model. For transcription, translation, or speech generation without a conversational agent, choose the dedicated API below.
28
29
30
31
53 32
5433```javascript| Build | Start here | What you control |
5534import { RealtimeAgent, RealtimeSession } from "@openai/agents/realtime";| ------------------------------------------------------------------ | ---------------------------------------------------------------------------- | ----------------------------------------------------------------- |
35| A speech-to-speech agent using the Realtime session and tool model | [Realtime API](https://developers.openai.com/api/docs/guides/realtime) | Audio turns, session state, tools, and interruptions. |
36| A voice interface for an existing text agent | [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents#build-a-chained-voice-workflow) | Speech-to-text, the text-agent workflow, then text-to-speech. |
37| A transcript of an audio file | [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text) | File uploads, bounded requests, and supported transcript formats. |
38| Live captions without assistant speech | [Live transcription](https://developers.openai.com/api/docs/guides/realtime-transcription) | Streaming audio and incremental transcript events. |
39| Continuous speech translation | [Live translation](https://developers.openai.com/api/docs/guides/realtime-translation) | A dedicated translation session, not a voice-agent turn loop. |
40| Narration or generated speech | [Text to speech](https://developers.openai.com/api/docs/guides/text-to-speech) | Text, voice, and output format. |
41| Audio input or output in an existing chat app | [Audio in Chat Completions](https://developers.openai.com/api/docs/guides/audio-chat-completions) | Bounded multimodal chat requests. |
56 42
5743const agent = new RealtimeAgent({## Build with voice
58 name: "Assistant",
59 instructions: "You are a helpful voice assistant.",
60});
61 44
6245const session = new RealtimeSession(agent, {Use [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) to compare architectures. Start with the prompting guide for [GPT-Live](https://developers.openai.com/api/docs/guides/live-prompting) or [Realtime](https://developers.openai.com/api/docs/guides/voice-prompting). Then use the shared guides for [custom voices](https://developers.openai.com/api/docs/guides/custom-voices), [evaluation](https://developers.openai.com/api/docs/guides/voice-agents#evaluate-your-voice-agent), and [cost optimization](https://developers.openai.com/api/docs/guides/voice-latency-cost). Each guide distinguishes model- or API-specific behavior.
6346 model: "gpt-realtime-2.1",
6447});## Choose a connection
48
49For browser audio, start with [WebRTC](https://developers.openai.com/api/docs/guides/voice-webrtc). For server audio pipelines, use [WebSockets](https://developers.openai.com/api/docs/guides/voice-websockets). For phone calls, see [Telephony and SIP](https://developers.openai.com/api/docs/guides/voice-sip). A [server-side control connection](https://developers.openai.com/api/docs/guides/voice-server-controls) lets a trusted backend observe and control a media session.
50
51Select your API on each connection page. Sharing a transport does not make GPT-Live and Realtime handshakes, credentials, or event formats interchangeable. Check the connection guide for prerequisites and setup instructions.
52
53## Add audio to your existing application
65 54
6655await session.connect({The Chat Completions examples now live in [Audio in Chat Completions](https://developers.openai.com/api/docs/guides/audio-chat-completions). For a browser voice-agent starter, use the [GPT-Live WebRTC quickstart](https://developers.openai.com/api/docs/guides/voice-webrtc?api=live).
67 apiKey: "ek_...(ephemeral key from your server)",
68});
69```
70
71
72This JavaScript example uses the Agents SDK to connect browser voice agents with WebRTC from the client. For Python voice workflows, use the [Voice agents guide](https://developers.openai.com/api/docs/guides/voice-agents), which covers chained voice pipelines.
73
74If you already have a text-based LLM application with the [Chat Completions endpoint](https://developers.openai.com/api/reference/resources/chat), you may want to add audio capabilities. For example, if your chat application supports text input, you can add audio input and output: include `audio` in the `modalities` array and use an audio model, like [`gpt-audio-1.5`](https://developers.openai.com/api/docs/models/gpt-audio-1.5).
75
76The [Responses API](https://developers.openai.com/api/reference/resources/responses) docs currently describe
77 text and image inputs with text outputs. For this audio-chat pattern, use Chat
78 Completions with an audio-capable model.
79
80
81
82Audio output from model
83
84 Create a human-like audio response to a prompt
85
86```javascript
87import { writeFileSync } from "node:fs";
88import OpenAI from "openai";
89
90const openai = new OpenAI();
91
92// Generate an audio response to the given prompt
93const response = await openai.chat.completions.create({
94 model: "gpt-audio-1.5",
95 modalities: ["text", "audio"],
96 audio: { voice: "alloy", format: "wav" },
97 messages: [
98 {
99 role: "user",
100 content: "Is a golden retriever a good family dog?",
101 },
102 ],
103 store: true,
104});
105
106// Inspect returned data
107console.log(response.choices[0]);
108
109// Write audio data to a file
110writeFileSync(
111 "dog.wav",
112 Buffer.from(response.choices[0].message.audio.data, "base64"),
113 { encoding: "utf-8" }
114);
115```
116
117```python
118import base64
119from openai import OpenAI
120
121client = OpenAI()
122
123completion = client.chat.completions.create(
124 model="gpt-audio-1.5",
125 modalities=["text", "audio"],
126 audio={"voice": "alloy", "format": "wav"},
127 messages=[{"role": "user", "content": "Is a golden retriever a good family dog?"}],
128)
129
130print(completion.choices[0])
131
132wav_bytes = base64.b64decode(completion.choices[0].message.audio.data)
133with open("dog.wav", "wb") as f:
134 f.write(wav_bytes)
135```
136
137```go
138package main
139
140import (
141 "context"
142 "encoding/base64"
143 "fmt"
144 "os"
145
146 "github.com/openai/openai-go/v3"
147)
148
149func main() {
150 client := openai.NewClient()
151 response, err := client.Chat.Completions.New(context.Background(), openai.ChatCompletionNewParams{
152 Model: "gpt-audio-1.5",
153 Modalities: []string{"text", "audio"},
154 Audio: openai.ChatCompletionAudioParam{
155 Voice: openai.ChatCompletionAudioParamVoiceUnion{OfString: openai.String("alloy")},
156 Format: openai.ChatCompletionAudioParamFormatWAV,
157 },
158 Messages: []openai.ChatCompletionMessageParamUnion{openai.UserMessage("Is a golden retriever a good family dog?")},
159 })
160 if err != nil {
161 panic(err)
162 }
163 fmt.Println(response.Choices[0])
164 audio, err := base64.StdEncoding.DecodeString(response.Choices[0].Message.Audio.Data)
165 if err != nil {
166 panic(err)
167 }
168 if err := os.WriteFile("dog.wav", audio, 0o600); err != nil {
169 panic(err)
170 }
171}
172```
173
174```java
175import com.openai.client.OpenAIClient;
176import com.openai.client.okhttp.OpenAIOkHttpClient;
177import com.openai.models.chat.completions.ChatCompletionAudioParam;
178import com.openai.models.chat.completions.ChatCompletionCreateParams;
179import java.io.IOException;
180import java.nio.file.Files;
181import java.nio.file.Path;
182import java.util.Base64;
183
184ChatCompletionCreateParams params =
185 ChatCompletionCreateParams.builder()
186 .model("gpt-audio-1.5")
187 .addUserMessage("Is a golden retriever a good family dog?")
188 .addModality(ChatCompletionCreateParams.Modality.TEXT)
189 .addModality(ChatCompletionCreateParams.Modality.AUDIO)
190 .audio(
191 ChatCompletionAudioParam.builder()
192 .voice("alloy")
193 .format(ChatCompletionAudioParam.Format.WAV)
194 .build())
195 .store(true)
196 .build();
197
198var message = client.chat().completions().create(params).choices().get(0).message();
199var audio =
200 message.audio().orElseThrow(() -> new IllegalStateException("No audio output returned"));
201Files.write(Path.of("dog.wav"), Base64.getDecoder().decode(audio.data()));
202message.content().ifPresent(System.out::println);
203```
204
205```csharp
206using OpenAI.Chat;
207#pragma warning disable OPENAI001
208
209string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;
210ChatClient client = new("gpt-audio-1.5", key);
211
212ChatCompletionOptions options = new()
213{
214 ResponseModalities = ChatResponseModalities.Text | ChatResponseModalities.Audio,
215 AudioOptions = new(ChatOutputAudioVoice.Alloy, ChatOutputAudioFormat.Wav),
216 StoredOutputEnabled = true,
217};
218
219ChatCompletion completion = await client.CompleteChatAsync(
220 [new UserChatMessage("Is a golden retriever a good family dog?")],
221 options
222);
223
224if (completion.OutputAudio is not ChatOutputAudio audio)
225{
226 throw new InvalidOperationException("No audio output was returned.");
227}
228
229Console.WriteLine(audio.Transcript);
230await File.WriteAllBytesAsync("dog.wav", audio.AudioBytes.ToArray());
231```
232
233```ruby
234require "base64"
235require "openai"
236
237client = OpenAI::Client.new
238completion = client.chat.completions.create(
239 model: "gpt-audio-1.5",
240 messages: [{role: :user, content: "Is a golden retriever a good family dog?"}],
241 modalities: [:text, :audio],
242 audio: {voice: :alloy, format: :wav},
243 store: true
244)
245
246audio = completion.choices.fetch(0).message.audio or raise "No audio returned"
247File.binwrite("dog.wav", Base64.strict_decode64(audio.data))
248```
249
250```bash
251curl "https://api.openai.com/v1/chat/completions" \
252 -H "Content-Type: application/json" \
253 -H "Authorization: Bearer $OPENAI_API_KEY" \
254 -d '{
255 "model": "gpt-audio-1.5",
256 "modalities": ["text", "audio"],
257 "audio": { "voice": "alloy", "format": "wav" },
258 "messages": [
259 {
260 "role": "user",
261 "content": "Is a golden retriever a good family dog?"
262 }
263 ]
264 }'
265```
266
267
268
269
270
271
272Audio input to model
273
274 Use audio inputs for prompting a model
275
276```javascript
277import OpenAI from "openai";
278const openai = new OpenAI();
279
280// Fetch an audio file and convert it to a base64 string
281const url = "https://cdn.openai.com/API/docs/audio/alloy.wav";
282const audioResponse = await fetch(url);
283const buffer = await audioResponse.arrayBuffer();
284const base64str = Buffer.from(buffer).toString("base64");
285
286const response = await openai.chat.completions.create({
287 model: "gpt-audio-1.5",
288 modalities: ["text", "audio"],
289 audio: { voice: "alloy", format: "wav" },
290 messages: [
291 {
292 role: "user",
293 content: [
294 { type: "text", text: "What is in this recording?" },
295 {
296 type: "input_audio",
297 input_audio: { data: base64str, format: "wav" },
298 },
299 ],
300 },
301 ],
302 store: true,
303});
304
305console.log(response.choices[0]);
306```
307
308```python
309import base64
310import requests
311from openai import OpenAI
312
313client = OpenAI()
314
315# Fetch the audio file and convert it to a base64 encoded string
316url = "https://cdn.openai.com/API/docs/audio/alloy.wav"
317response = requests.get(url)
318response.raise_for_status()
319wav_data = response.content
320encoded_string = base64.b64encode(wav_data).decode("utf-8")
321
322completion = client.chat.completions.create(
323 model="gpt-audio-1.5",
324 modalities=["text", "audio"],
325 audio={"voice": "alloy", "format": "wav"},
326 messages=[
327 {
328 "role": "user",
329 "content": [
330 {"type": "text", "text": "What is in this recording?"},
331 {
332 "type": "input_audio",
333 "input_audio": {"data": encoded_string, "format": "wav"},
334 },
335 ],
336 },
337 ],
338)
339
340print(completion.choices[0].message)
341```
342
343```go
344package main
345
346import (
347 "context"
348 "encoding/base64"
349 "fmt"
350 "os"
351
352 "github.com/openai/openai-go/v3"
353)
354
355func main() {
356 audio, err := os.ReadFile("fixtures/audio.wav")
357 if err != nil {
358 panic(err)
359 }
360 client := openai.NewClient()
361 response, err := client.Chat.Completions.New(context.Background(), openai.ChatCompletionNewParams{
362 Model: "gpt-audio-1.5",
363 Modalities: []string{"text", "audio"},
364 Audio: openai.ChatCompletionAudioParam{
365 Voice: openai.ChatCompletionAudioParamVoiceUnion{OfString: openai.String("alloy")},
366 Format: openai.ChatCompletionAudioParamFormatWAV,
367 },
368 Messages: []openai.ChatCompletionMessageParamUnion{openai.UserMessage([]openai.ChatCompletionContentPartUnionParam{
369 openai.TextContentPart("What is in this recording?"),
370 openai.InputAudioContentPart(openai.ChatCompletionContentPartInputAudioInputAudioParam{
371 Data: base64.StdEncoding.EncodeToString(audio),
372 Format: "wav",
373 }),
374 })},
375 })
376 if err != nil {
377 panic(err)
378 }
379 fmt.Println(response.Choices[0])
380}
381```
382
383```java
384import com.openai.client.OpenAIClient;
385import com.openai.client.okhttp.OpenAIOkHttpClient;
386import com.openai.models.chat.completions.ChatCompletionAudioParam;
387import com.openai.models.chat.completions.ChatCompletionContentPart;
388import com.openai.models.chat.completions.ChatCompletionContentPartInputAudio;
389import com.openai.models.chat.completions.ChatCompletionContentPartText;
390import com.openai.models.chat.completions.ChatCompletionCreateParams;
391import com.openai.models.chat.completions.ChatCompletionUserMessageParam;
392import java.io.IOException;
393import java.nio.file.Files;
394import java.nio.file.Path;
395import java.util.Base64;
396import java.util.List;
397
398String encodedAudio =
399 Base64.getEncoder()
400 .encodeToString(
401 Files.readAllBytes(Path.of(System.getenv("OPENAI_EXAMPLE_AUDIO_PATH"))));
402
403ChatCompletionCreateParams params =
404 ChatCompletionCreateParams.builder()
405 .model("gpt-audio-1.5")
406 .addMessage(
407 ChatCompletionUserMessageParam.builder()
408 .contentOfArrayOfContentParts(
409 List.of(
410 ChatCompletionContentPart.ofText(
411 ChatCompletionContentPartText.builder()
412 .text("What is in this recording?")
413 .build()),
414 ChatCompletionContentPart.ofInputAudio(
415 ChatCompletionContentPartInputAudio.builder()
416 .inputAudio(
417 ChatCompletionContentPartInputAudio.InputAudio.builder()
418 .data(encodedAudio)
419 .format(
420 ChatCompletionContentPartInputAudio.InputAudio
421 .Format.WAV)
422 .build())
423 .build())))
424 .build())
425 .addModality(ChatCompletionCreateParams.Modality.TEXT)
426 .addModality(ChatCompletionCreateParams.Modality.AUDIO)
427 .audio(
428 ChatCompletionAudioParam.builder()
429 .voice("alloy")
430 .format(ChatCompletionAudioParam.Format.WAV)
431 .build())
432 .store(true)
433 .build();
434
435client.chat().completions().create(params).choices().stream()
436 .flatMap(choice -> choice.message().content().stream())
437 .forEach(System.out::println);
438```
439
440```csharp
441using OpenAI.Chat;
442#pragma warning disable OPENAI001
443
444string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;
445ChatClient client = new("gpt-audio-1.5", key);
446
447BinaryData audio = BinaryData.FromBytes(
448 await File.ReadAllBytesAsync("audio.wav")
449);
450UserChatMessage message = new(
451 [
452 ChatMessageContentPart.CreateTextPart("What is in this recording?"),
453 ChatMessageContentPart.CreateInputAudioPart(
454 audio,
455 ChatInputAudioFormat.Wav
456 ),
457 ]
458);
459ChatCompletionOptions options = new()
460{
461 ResponseModalities = ChatResponseModalities.Text | ChatResponseModalities.Audio,
462 AudioOptions = new(ChatOutputAudioVoice.Alloy, ChatOutputAudioFormat.Wav),
463 StoredOutputEnabled = true,
464};
465
466ChatCompletion completion = await client.CompleteChatAsync([message], options);
467
468if (completion.OutputAudio is not ChatOutputAudio audioOutput)
469{
470 throw new InvalidOperationException("No audio output was returned.");
471}
472
473Console.WriteLine(audioOutput.Transcript);
474```
475
476```ruby
477require "base64"
478require "openai"
479
480client = OpenAI::Client.new
481audio = Base64.strict_encode64(File.binread("audio.wav"))
482completion = client.chat.completions.create(
483 model: "gpt-audio-1.5",
484 messages: [{
485 role: :user,
486 content: [
487 {type: :text, text: "What is in this recording?"},
488 {type: :input_audio, input_audio: {data: audio, format: :wav}}
489 ]
490 }],
491 modalities: [:text, :audio],
492 audio: {voice: :alloy, format: :wav},
493 store: true
494)
495
496puts(completion.choices.fetch(0).message.content)
497```
498
499```bash
500curl "https://api.openai.com/v1/chat/completions" \
501 -H "Content-Type: application/json" \
502 -H "Authorization: Bearer $OPENAI_API_KEY" \
503 -d '{
504 "model": "gpt-audio-1.5",
505 "modalities": ["text", "audio"],
506 "audio": { "voice": "alloy", "format": "wav" },
507 "messages": [
508 {
509 "role": "user",
510 "content": [
511 { "type": "text", "text": "What is in this recording?" },
512 {
513 "type": "input_audio",
514 "input_audio": {
515 "data": "<base64 bytes here>",
516 "format": "wav"
517 }
518 }
519 ]
520 }
521 ]
522 }'
523```