1# Voice agents1# Voice agents
2 2
3import {
4 Bolt,
5 Cube,
6 Desktop,
7 Phone,
8} from "@components/react/oai/platform/ui/Icon.react";
9
10
3Voice agents turn the same agent concepts into spoken, low-latency interactions. The key design choice is deciding whether the model should work directly with live audio or whether your application should explicitly chain speech-to-text, text reasoning, and text-to-speech.11Voice agents turn the same agent concepts into spoken, low-latency interactions. The key design choice is deciding whether the model should work directly with live audio or whether your application should explicitly chain speech-to-text, text reasoning, and text-to-speech.
4 12
5## Choose the right architecture13## Choose the right architecture
13 21
14## Recommended starting points22## Recommended starting points
15 23
16The two supported languages expose different strengths today:24The examples below are intentionally different architectures, not matching language tabs. The TypeScript and Python libraries expose different voice helpers today:
17 25
18- In TypeScript, the fastest path to a browser-based voice assistant is a `RealtimeAgent` and `RealtimeSession`.26- In TypeScript, the fastest path to a browser-based voice assistant is a `RealtimeAgent` and `RealtimeSession`.
19- In Python, the simplest path to extending an existing text agent into voice is a chained `VoicePipeline`.27- In Python, the simplest path to extending an existing text agent into voice is a chained `VoicePipeline`.
20 28
21Two common voice starting points
22
23```typescript
24import { RealtimeAgent, RealtimeSession } from "@openai/agents/realtime";
25
26const agent = new RealtimeAgent({
27 name: "Assistant",
28 instructions: "You are a helpful voice assistant.",
29});
30
31const session = new RealtimeSession(agent, {
32 model: "gpt-realtime-1.5",
33});
34
35await session.connect({
36 apiKey: "ek_...(ephemeral key from your server)",
37});
38```
39
40```python
41import asyncio
42import numpy as np
43
44from agents import Agent, function_tool
45from agents.voice import AudioInput, SingleAgentVoiceWorkflow, VoicePipeline
46
47
48@function_tool
49def get_weather(city: str) -> str:
50 """Get the weather for a given city."""
51 return f"The weather in {city} is sunny."
52
53
54agent = Agent(
55 name="Assistant",
56 instructions="You are a helpful voice assistant.",
57 model="gpt-5.5",
58 tools=[get_weather],
59)
60
61
62async def main() -> None:
63 pipeline = VoicePipeline(workflow=SingleAgentVoiceWorkflow(agent))
64 audio_input = AudioInput(buffer=np.zeros(24000 * 3, dtype=np.int16))
65 result = await pipeline.run(audio_input)
66 async for event in result.stream():
67 if event.type == "voice_stream_event_audio":
68 print("Received audio bytes", len(event.data))
69
70
71if __name__ == "__main__":
72 asyncio.run(main())
73```
74
75
76<span id="speech-to-speech-realtime-architecture"></span>29<span id="speech-to-speech-realtime-architecture"></span>
77 30
78## Build a speech-to-speech voice agent31## Build a speech-to-speech voice agent
79 32
80Use the live audio API path when the interaction should feel conversational and immediate. The usual browser flow is:33Use the live audio API path when the interaction should feel conversational and immediate. This is the best starting point for voice agents that need barge-in, low first-audio latency, natural turn taking, and realtime tool use.
34
35The usual browser flow is:
81 36
821. Your application server creates an ephemeral client secret for the live audio session.371. Your application server creates an ephemeral client secret for the live audio session.
832. Your frontend creates a `RealtimeSession`.382. Your frontend creates a `RealtimeSession`.
843. The session connects over WebRTC in the browser or WebSocket on the server.393. The session connects over WebRTC in the browser or WebSocket on the server.
854. The agent handles audio turns, tools, interruptions, and handoffs inside that session.404. The agent handles audio turns, tools, interruptions, and handoffs inside that session.
86 41
42From there, attach tools, handoffs, and guardrails to the `RealtimeAgent` the same way you would attach them to a text agent. Keep audio transport concerns in the session layer, and keep business logic in the agent definition.
43
87Start with the transport docs when you need lower-level control:44Start with the transport docs when you need lower-level control:
88 45
89- [Live audio API overview](https://developers.openai.com/api/docs/guides/realtime)46- [Realtime and audio overview](https://developers.openai.com/api/docs/guides/realtime)
90- [Live audio API with WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc)47- [Live audio API with WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc)
91- [Live audio API with WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket)48- [Live audio API with WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket)
92 49
100 57
101This is often the better fit for support flows, approval-heavy flows, or cases where you want durable transcripts and deterministic logic between each stage.58This is often the better fit for support flows, approval-heavy flows, or cases where you want durable transcripts and deterministic logic between each stage.
102 59
60Use this path when each stage needs to be visible or replaceable. For example, you might store the transcript, run policy checks before the text agent responds, call internal systems, then generate speech only after the workflow reaches an approved answer.
61
103## Voice agents still use the same core agent building blocks62## Voice agents still use the same core agent building blocks
104 63
105The voice surface changes the transport and audio loop, but the core workflow decisions are the same:64The voice surface changes the transport and audio loop, but the core workflow decisions are the same:
111- Use [Integrations and observability](https://developers.openai.com/api/docs/guides/agents/integrations-observability) when you need MCP-backed capabilities or want to inspect how the voice workflow behaved.70- Use [Integrations and observability](https://developers.openai.com/api/docs/guides/agents/integrations-observability) when you need MCP-backed capabilities or want to inspect how the voice workflow behaved.
112 71
113The practical rule is: choose the audio architecture first, then design the rest of the agent workflow the same way you would for text.72The practical rule is: choose the audio architecture first, then design the rest of the agent workflow the same way you would for text.
73
74## Next steps
75
76<a href="/api/docs/guides/realtime">
77
78
79<span slot="icon">
80 </span>
81 Choose the right realtime or audio guide for your use case.
82
83
84</a>
85
86<a href="/api/docs/guides/realtime-conversations">
87
88
89<span slot="icon">
90 </span>
91 Work with the Realtime session lifecycle and event model.
92
93
94</a>
95
96<a href="/api/docs/guides/realtime-webrtc">
97
98
99<span slot="icon">
100 </span>
101 Connect browser and mobile audio directly to a Realtime session.
102
103
104</a>
105
106<a href="/api/docs/guides/realtime-models-prompting">
107
108
109<span slot="icon">
110 </span>
111 Tune reasoning, preambles, tools, entity capture, and voice behavior.
112
113
114</a>