SpyBara
Go Premium

Documentation 2026-09-11 20:00 UTC to 2026-09-13 15:02 UTC

45 files changed +275 −477. View all changes and history on the product overview
2026
Tue 29 22:57 Mon 28 22:57 Sat 26 23:59 Fri 25 23:58 Thu 24 23:58 Wed 23 23:58 Tue 22 23:57 Mon 21 23:00 Sat 19 23:00 Fri 18 22:59 Thu 17 10:04 Wed 16 20:58 Tue 15 22:59 Mon 14 22:58 Sun 13 15:02 Fri 11 20:00 Thu 10 18:01 Wed 9 23:59 Sat 5 17:01 Fri 4 23:59 Thu 3 23:00 Wed 2 22:59

deprecations.md +28 −20

Details

34 34 

35Upcoming deprecations are listed below, with the most recent announcements at the top.35Upcoming deprecations are listed below, with the most recent announcements at the top.

36 36 

37### 2026-09-11: GPT-5.4-Cyber

38 

39The `gpt-5.4-cyber` model is deprecated and will be removed from the API on October 1, 2026. Migrate to `gpt-5.6-cyber` before the shutdown date.

40 

41| Shutdown date | Model / system | Recommended replacement |

42| ------------- | --------------- | ----------------------- |

43| Oct 1, 2026 | `gpt-5.4-cyber` | `gpt-5.6-cyber` |

44 

37### 2026-08-26: Transcription models45### 2026-08-26: Transcription models

38 46 

39On August 26, 2026, we notified developers using `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, and `gpt-4o-transcribe-diarize` of their deprecation and removal from the API on February 26, 2027.47On August 26, 2026, we notified developers using `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, and `gpt-4o-transcribe-diarize` of their deprecation and removal from the API on February 26, 2027.


191 199 

192Past deprecations are listed below, with the most recent announcements at the top.200Past deprecations are listed below, with the most recent announcements at the top.

193 201 

194### 2026-05-08: gpt-5.2-chat-latest and gpt-5.3-chat-latest model snapshots202### 2026-05-08: `gpt-5.2-chat-latest` and `gpt-5.3-chat-latest` model snapshots

195 203 

196On May 8th, 2026, we notified developers using `gpt-5.2-chat-latest` and `gpt-5.3-chat-latest` model snapshots of their deprecation and removal from the API.204On May 8th, 2026, we notified developers using `gpt-5.2-chat-latest` and `gpt-5.3-chat-latest` model snapshots of their deprecation and removal from the API.

197 205 


221| July 23, 2026 | `o4-mini-deep-research-2025-06-26` \| `o4-mini-deep-research` | `gpt-5.6-sol` |229| July 23, 2026 | `o4-mini-deep-research-2025-06-26` \| `o4-mini-deep-research` | `gpt-5.6-sol` |

222| July 23, 2026 | `gpt-5.2-codex` | `gpt-5.6-sol` |230| July 23, 2026 | `gpt-5.2-codex` | `gpt-5.6-sol` |

223 231 

224### 2025-11-18: chatgpt-4o-latest snapshot232### 2025-11-18: `chatgpt-4o-latest` snapshot

225 233 

226On November 18th, 2025, we notified developers using `chatgpt-4o-latest` model snapshot of its deprecation and removal from the API on February 17, 2026.234On November 18th, 2025, we notified developers using `chatgpt-4o-latest` model snapshot of its deprecation and removal from the API on February 17, 2026.

227 235 


229| ------------- | ------------------- | ----------------------- |237| ------------- | ------------------- | ----------------------- |

230| 2026-02-17 | `chatgpt-4o-latest` | `gpt-5.1-chat-latest` |238| 2026-02-17 | `chatgpt-4o-latest` | `gpt-5.1-chat-latest` |

231 239 

232### 2025-11-17: codex-mini-latest model snapshot240### 2025-11-17: `codex-mini-latest` model snapshot

233 241 

234On November 17th, 2025, we notified developers using `codex-mini-latest` model of its deprecation and removal from the API on February 12, 2026. As part of this deprecation, we will no longer support our legacy local shell tool, which is only available for use with `codex-mini-latest`. For new use cases, please use our latest shell tool.242On November 17th, 2025, we notified developers using `codex-mini-latest` model of its deprecation and removal from the API on February 12, 2026. As part of this deprecation, we will no longer support our legacy local shell tool, which is only available for use with `codex-mini-latest`. For new use cases, please use our latest shell tool.

235 243 


262 270 

263The Realtime API Beta was deprecated and removed from the API on May 12, 2026.271The Realtime API Beta was deprecated and removed from the API on May 12, 2026.

264 272 

265There are a few key differences between the interfaces in the Realtime beta API and the released GA API. See [the migration guide](https://developers.openai.com/api/docs/guides/realtime#beta-to-ga-migration) for the current GA interface and related Realtime docs.273The interfaces in the Realtime beta API and the released GA API have a few key differences. See [the migration guide](https://developers.openai.com/api/docs/guides/realtime#beta-to-ga-migration) for the current GA interface and related Realtime docs.

266 274 

267| Shutdown date | Model / system | Recommended replacement |275| Shutdown date | Model / system | Recommended replacement |

268| ------------- | ------------------------ | ----------------------- |276| ------------- | ------------------------ | ----------------------- |

269| 2026‑05‑12 | OpenAI-Beta: realtime=v1 | Realtime API |277| 2026‑05‑12 | OpenAI-Beta: realtime=v1 | Realtime API |

270 278 

271### 2025-09-15: gpt-4o-realtime-preview models279### 2025-09-15: `gpt-4o-realtime-preview` models

272 280 

273In September, 2025, we notified developers using gpt-4o-realtime-preview models of their deprecation and removal from the API in six months.281In September, 2025, we notified developers using `gpt-4o-realtime-preview` models of their deprecation and removal from the API in six months.

274 282 

275| Shutdown date | Model / system | Recommended replacement |283| Shutdown date | Model / system | Recommended replacement |

276| ------------- | ---------------------------------- | ----------------------- |284| ------------- | ------------------------------------ | ----------------------- |

277| 2026-05-07 | gpt-4o-realtime-preview | gpt-realtime-1.5 |285| 2026-05-07 | `gpt-4o-realtime-preview` | `gpt-realtime-1.5` |

278| 2026-05-07 | gpt-4o-realtime-preview-2025-06-03 | gpt-realtime-1.5 |286| 2026-05-07 | `gpt-4o-realtime-preview-2025-06-03` | `gpt-realtime-1.5` |

279| 2026-05-07 | gpt-4o-realtime-preview-2024-12-17 | gpt-realtime-1.5 |287| 2026-05-07 | `gpt-4o-realtime-preview-2024-12-17` | `gpt-realtime-1.5` |

280| 2026-05-07 | gpt-4o-mini-realtime-preview | gpt-realtime-mini |288| 2026-05-07 | `gpt-4o-mini-realtime-preview` | `gpt-realtime-mini` |

281| 2026-05-07 | gpt-4o-audio-preview | gpt-audio-1.5 |289| 2026-05-07 | `gpt-4o-audio-preview` | `gpt-audio-1.5` |

282| 2026-05-07 | gpt-4o-mini-audio-preview | gpt-audio-mini |290| 2026-05-07 | `gpt-4o-mini-audio-preview` | `gpt-audio-mini` |

283 291 

284### 2025-08-20: Assistants API292### 2025-08-20: Assistants API

285 293 


293| ------------- | -------------- | ----------------------------------- |301| ------------- | -------------- | ----------------------------------- |

294| 2026‑08‑26 | Assistants API | Responses API and Conversations API |302| 2026‑08‑26 | Assistants API | Responses API and Conversations API |

295 303 

296### 2025-06-10: gpt-4o-realtime-preview-2024-10-01304### 2025-06-10: `gpt-4o-realtime-preview-2024-10-01`

297 305 

298On June 10th, 2025, we notified developers using gpt-4o-realtime-preview-2024-10-01 of its deprecation and removal from the API in three months.306On June 10th, 2025, we notified developers using `gpt-4o-realtime-preview-2024-10-01` of its deprecation and removal from the API in three months.

299 307 

300| Shutdown date | Model / system | Recommended replacement |308| Shutdown date | Model / system | Recommended replacement |

301| ------------- | ---------------------------------- | ----------------------- |309| ------------- | ------------------------------------ | ----------------------- |

302| 2025-10-10 | gpt-4o-realtime-preview-2024-10-01 | gpt-realtime-1.5 |310| 2025-10-10 | `gpt-4o-realtime-preview-2024-10-01` | `gpt-realtime-1.5` |

303 311 

304### 2025-06-10: gpt-4o-audio-preview-2024-10-01312### 2025-06-10: `gpt-4o-audio-preview-2024-10-01`

305 313 

306On June 10th, 2025, we notified developers using `gpt-4o-audio-preview-2024-10-01` of its deprecation and removal from the API in three months.314On June 10th, 2025, we notified developers using `gpt-4o-audio-preview-2024-10-01` of its deprecation and removal from the API in three months.

307 315 


309| ------------- | --------------------------------- | ----------------------- |317| ------------- | --------------------------------- | ----------------------- |

310| 2025-10-10 | `gpt-4o-audio-preview-2024-10-01` | `gpt-audio-1.5` |318| 2025-10-10 | `gpt-4o-audio-preview-2024-10-01` | `gpt-audio-1.5` |

311 319 

312### 2025-04-28: text-moderation320### 2025-04-28: `text-moderation`

313 321 

314On April 28th, 2025, we notified developers using `text-moderation` of its deprecation and removal from the API in six months.322On April 28th, 2025, we notified developers using `text-moderation` of its deprecation and removal from the API in six months.

315 323 


319| 2025-10-27 | `text-moderation-stable` | `omni-moderation` |327| 2025-10-27 | `text-moderation-stable` | `omni-moderation` |

320| 2025-10-27 | `text-moderation-latest` | `omni-moderation` |328| 2025-10-27 | `text-moderation-latest` | `omni-moderation` |

321 329 

322### 2025-04-28: o1-preview and o1-mini330### 2025-04-28: `o1-preview` and `o1-mini`

323 331 

324On April 28th, 2025, we notified developers using `o1-preview` and `o1-mini` of their deprecations and removal from the API in three months and six months respectively.332On April 28th, 2025, we notified developers using `o1-preview` and `o1-mini` of their deprecations and removal from the API in three months and six months respectively.

325 333 

Details

359// Replace the illustrative IDs and URLs below with your own resource values.359// Replace the illustrative IDs and URLs below with your own resource values.

360import OpenAI from "openai";360import OpenAI from "openai";

361 361 

362/**

363 * @param {OpenAI} client

364 * @param {string} sessionId

365 */

366async function deleteSession(client, sessionId) {362async function deleteSession(client, sessionId) {

367 return client.beta.agents.sessions.delete(sessionId);363 return client.beta.agents.sessions.delete(sessionId);

368}364}

Details

185// Replace the illustrative IDs and URLs below with your own resource values.185// Replace the illustrative IDs and URLs below with your own resource values.

186import OpenAI from "openai";186import OpenAI from "openai";

187 187 

188/**

189 * @param {OpenAI} client

190 * @param {string} sessionId

191 */

192async function deleteSession(client, sessionId) {188async function deleteSession(client, sessionId) {

193 return client.beta.agents.sessions.delete(sessionId);189 return client.beta.agents.sessions.delete(sessionId);

194}190}

Details

32const client = new OpenAI();32const client = new OpenAI();

33const model = "gpt-6-astra";33const model = "gpt-6-astra";

34 34 

35/** @type {OpenAI.Responses.FunctionTool[]} */

36const tools = [35const tools = [

37 {36 {

38 type: "function",37 type: "function",

Details

55 thread.id,55 thread.id,

56 HiddenContextItem(56 HiddenContextItem(

57 id="item_123",57 id="item_123",

58 thread_id=thread.id,

58 created_at=datetime.now(),59 created_at=datetime.now(),

59 content="<USER_ACTION>The user did a thing</USER_ACTION>",60 content="<USER_ACTION>The user did a thing</USER_ACTION>",

60 ),61 ),

Details

232const SOURCE_ID_RE = /^[A-Za-z0-9_-]+$/;232const SOURCE_ID_RE = /^[A-Za-z0-9_-]+$/;

233const LINE_LOCATOR_RE = /^L\d+(?:-L\d+)?$/;233const LINE_LOCATOR_RE = /^L\d+(?:-L\d+)?$/;

234 234 

235/**235// Extract citations such as:

236 * @typedef {Object} Citation236//

237 * @property {string} raw237// {CITATION_START}cite{CITATION_DELIMITER}turn0file0{CITATION_STOP}

238 * @property {string} family238// {CITATION_START}cite{CITATION_DELIMITER}turn0file0{CITATION_DELIMITER}L8-L13{CITATION_STOP}

239 * @property {string[]} source_ids239// {CITATION_START}cite{CITATION_DELIMITER}turn0search0{CITATION_DELIMITER}turn1news2{CITATION_STOP}

240 * @property {string | null} locator

241 * @property {number} start

242 * @property {number} end

243 */

244 

245/**

246 * Extract citations such as:

247 *

248 * {CITATION_START}cite{CITATION_DELIMITER}turn0file0{CITATION_STOP}

249 * {CITATION_START}cite{CITATION_DELIMITER}turn0file0{CITATION_DELIMITER}L8-L13{CITATION_STOP}

250 * {CITATION_START}cite{CITATION_DELIMITER}turn0search0{CITATION_DELIMITER}turn1news2{CITATION_STOP}

251 *

252 * @param {string} text

253 * @param {{ families?: string[] }} [options]

254 * @returns {Citation[]}

255 */

256function extractCitations(text, { families = ["cite"] } = {}) {240function extractCitations(text, { families = ["cite"] } = {}) {

257 if (families.length === 0) {241 if (families.length === 0) {

258 return [];242 return [];


267 "g"251 "g"

268 );252 );

269 253 

270 /** @type {Citation[]} */

271 const citations = [];254 const citations = [];

272 255 

273 for (const match of text.matchAll(tokenRe)) {256 for (const match of text.matchAll(tokenRe)) {


304 return citations;287 return citations;

305}288}

306 289 

307/**

308 * @param {string} text

309 * @param {Iterable<Citation>} citations

310 * @returns {string}

311 */

312function stripCitations(text, citations) {290function stripCitations(text, citations) {

313 let cleanText = text;291 let cleanText = text;

314 const sortedCitations = Array.from(citations).sort(292 const sortedCitations = Array.from(citations).sort(


326 304 

327```python305```python

328import re306import re

329from typing import Iterable, TypedDict307from collections.abc import Iterable

308from typing import TypedDict

330 309 

331CITATION_START = "\ue200"310CITATION_START = "\ue200"

332CITATION_DELIMITER = "\ue202"311CITATION_DELIMITER = "\ue202"

Details

58 58 

59const client = new OpenAI();59const client = new OpenAI();

60 60 

61/** @type {import("openai/resources/responses/responses").ResponseInput} */

62const conversation = [61const conversation = [

63 {62 {

64 type: "message",63 type: "message",


302 301 

303const client = new OpenAI();302const client = new OpenAI();

304 303 

305/** @type {import("openai/resources/responses/responses").ResponseInput} */

306const conversation = [{ role: "user", content: "Plan a trip to Kyoto." }];304const conversation = [{ role: "user", content: "Plan a trip to Kyoto." }];

307 305 

308const compacted = await client.responses.compact({306const compacted = await client.responses.compact({


310 input: conversation,308 input: conversation,

311});309});

312 310 

313/** @type {import("openai/resources/responses/responses").ResponseInput} */

314const nextInput = [311const nextInput = [

315 ...compacted.output.map(312 ...compacted.output.map((item) => item),

316 (item) =>

317 /** @type {import("openai/resources/responses/responses").ResponseInputItem} */ (

318 item

319 )

320 ),

321 { role: "user", content: "Add two more days to the itinerary." },313 { role: "user", content: "Add two more days to the itinerary." },

322];314];

323 315 

Details

188 188 

189const openai = new OpenAI();189const openai = new OpenAI();

190 190 

191/** @type {OpenAI.Responses.ResponseInput} */

192let history = [191let history = [

193 {192 {

194 role: "user",193 role: "user",

Details

18| [Use `reasoning.encrypted_content`](#use-reasoningencryptedcontent) | Quality, latency |18| [Use `reasoning.encrypted_content`](#use-reasoningencryptedcontent) | Quality, latency |

19| [Set image detail intentionally](#set-image-detail-intentionally) | Quality, cost, latency |19| [Set image detail intentionally](#set-image-detail-intentionally) | Quality, cost, latency |

20| [Send a safety identifier](#send-a-safety-identifier) | Safety, reliability |20| [Send a safety identifier](#send-a-safety-identifier) | Safety, reliability |

21| [Use `background=True`](#use-backgroundtrue) | Resumability |21| [Use `background=True`](#use-backgroundtrue) | Resuming work |

22| [Use WebSocket mode](#use-websocket-mode) | Latency |22| [Use WebSocket mode](#use-websocket-mode) | Latency |

23 23 

24## Use the Responses API24## Use the Responses API


51debugging, synthesis, and multi-step tradeoffs.51debugging, synthesis, and multi-step tradeoffs.

52 52 

53Use `low` when the job is mostly extraction, routing, classification, or a53Use `low` when the job is mostly extraction, routing, classification, or a

54simple rewrite. Use `medium` or `high` when the model needs to diagnose a54routine rewrite. Use `medium` or `high` when the model needs to diagnose a

55problem, compare options, write a plan, or reason through code. Use `xhigh` or55problem, compare options, write a plan, or reason through code. Use `xhigh` or

56`max` only when representative evals show that the quality gain justifies the56`max` only when representative evals show that the quality gain justifies the

57extra latency and cost. When migrating from GPT-5.5 or GPT-5.4, start with the57extra latency and cost. When migrating from GPT-5.5 or GPT-5.4, start with the


392are the deferred tool definitions loaded into context. Only then will the model392are the deferred tool definitions loaded into context. Only then will the model

393call them. This saves tokens and preserves cache performance.393call them. This saves tokens and preserves cache performance.

394 394 

395There are two modes:395Tool search has two modes:

396 396 

397- **Hosted tool search** is the simpler option. Use it when you already know397- **Hosted tool search** is the simpler option. Use it when you already know

398 which tools could be available for the request.398 which tools could be available for the request.


419 419 

420const openai = new OpenAI();420const openai = new OpenAI();

421 421 

422/** @type {OpenAI.Responses.Tool} */

423const billingNamespace = {422const billingNamespace = {

424 type: "namespace",423 type: "namespace",

425 name: "billing",424 name: "billing",


444 ],443 ],

445};444};

446 445 

447/** @type {OpenAI.Responses.Tool} */

448const crmNamespace = {446const crmNamespace = {

449 type: "namespace",447 type: "namespace",

450 name: "crm",448 name: "crm",


773 771 

774## Leverage built-in tools772## Leverage built-in tools

775 773 

776[Built-in tools](https://developers.openai.com/api/docs/guides/tools) are the API's native capabilities.774[Built-in tools](https://developers.openai.com/api/docs/guides/tools) are native capabilities of the API.

777Instead of building every tool yourself, you can give the model access to tools775Instead of building every tool yourself, you can give the model access to tools

778that already work inside the Responses API. The model can then decide when to776that already work inside the Responses API. The model can then decide when to

779use them.777use them.


794- **Skills**: Attach reusable instruction bundles and workflow files792- **Skills**: Attach reusable instruction bundles and workflow files

795- **Apply patch**: Make structured code edits793- **Apply patch**: Make structured code edits

796 794 

797There is also a model-quality reason to prefer them. Built-in tools are795Model quality is another reason to prefer them. Built-in tools are

798in-distribution for our post-training, meaning that the models are trained and796in-distribution for our post-training, meaning that the models are trained and

799evaluated around these tool shapes, behaviors, and outputs. With built-in tools,797evaluated around these tool shapes, behaviors, and outputs. With built-in tools,

800OpenAI models support better tool selection, cleaner execution, and fewer798OpenAI models support better tool selection, cleaner execution, and fewer


815next turn is built around the important state, not every intermediate reasoning,813next turn is built around the important state, not every intermediate reasoning,

816failed command, and obsolete branch of reasoning.814failed command, and obsolete branch of reasoning.

817 815 

818There are two ways to leverage compaction:816You can use compaction in two ways:

819 817 

820- **Let the server handle it**: if you use `previous_response_id`, turn on818- **Let the server handle it**: if you use `previous_response_id`, turn on

821 `context_management` with a `compact_threshold`. The server will automatically819 `context_management` with a `compact_threshold`. The server will automatically


1064compare write volume with later cache reads to measure net cost and tune1062compare write volume with later cache reads to measure net cost and tune

1065breakpoint placement.1063breakpoint placement.

1066 1064 

1067Use an optional `prompt_cache_key` to maintain separate cache accounting for1065Use a stable `prompt_cache_key` for requests that share a reusable prefix to

1068customers, users, or workspaces. This can make cached token usage and billing1066help route related requests to the same cache and optimize cache hit rates on

1067models before GPT-5.6. For busy groups, follow the [guidance for distributing

1068traffic across more keys](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-keys).

1069 

1070On GPT-5.6 and later, `prompt_cache_key` is optional: you can achieve optimal

1071cache hit rates without it. You can use it to maintain separate cache accounting

1072for customers, users, or workspaces. This can make cached token usage and billing

1069easier to explain for each group. Assign a distinct key to each customer and1073easier to explain for each group. Assign a distinct key to each customer and

1070keep it stable across that customer's related requests. Separate keys also help1074keep it stable across that customer's related requests. Separate keys also help

1071prevent cache-hit probing across customers. See [Separate cache accounting with1075prevent cache-hit probing across customers. See [Separate cache accounting with


1243 1247 

1244const openai = new OpenAI();1248const openai = new OpenAI();

1245 1249 

1246/** @type {OpenAI.Responses.ResponseInput} */

1247const history = [1250const history = [

1248 {1251 {

1249 role: "user",1252 role: "user",


1694[WebSocket mode](https://developers.openai.com/api/docs/guides/websocket-mode) is built for long-running,1697[WebSocket mode](https://developers.openai.com/api/docs/guides/websocket-mode) is built for long-running,

1695tool-call-heavy workflows where you keep a persistent connection open and1698tool-call-heavy workflows where you keep a persistent connection open and

1696continue by sending only new input items plus `previous_response_id`. For1699continue by sending only new input items plus `previous_response_id`. For

1697rollouts with 20 or more tool calls, this approach is roughly 40% faster1700runs with 20 or more tool calls, this approach is roughly 40% faster

1698end-to-end.1701end-to-end.

1699 1702 

1700**How this works**: The first message will look like a normal Responses request:1703**How this works**: The first message will look like a normal Responses request:


1720 1723 

1721The Python sample uses `pip install "openai[realtime]>=3.8.0"`.1724The Python sample uses `pip install "openai[realtime]>=3.8.0"`.

1722The JavaScript sample uses `npm install openai@^7.10.0 ws`.1725The JavaScript sample uses `npm install openai@^7.10.0 ws`.

1723The Ruby sample uses `gem install async-websocket`.1726The Ruby sample uses `gem install openai async-websocket`.

1724 1727 

1725Start a Responses API WebSocket session1728Start a Responses API WebSocket session

1726 1729 


1803 1806 

1804```ruby1807```ruby

1805require "async"1808require "async"

1806require "async/http/endpoint"1809require "openai"

1807require "async/websocket/client"

1808require "json"1810require "json"

1809 1811 

1810def wait_for_response(connection)1812def wait_for_response(connection)

1811 while (message = connection.read)1813 while (event = connection.receive)

1812 event = JSON.parse(message.to_str)1814 case event.type.to_s

1813 case event.fetch("type")1815 when "response.completed" then return event.response

1814 when "response.completed" then return event.fetch("response")

1815 when "response.failed", "response.incomplete", "error"1816 when "response.failed", "response.incomplete", "error"

1816 raise "Response failed: #{JSON.generate(event)}"1817 raise "Response failed: #{event.to_json}"

1817 end1818 end

1818 end1819 end

1819 raise "Connection closed before the response finished"1820 raise "Connection closed before the response finished"


1844 strict: true1845 strict: true

1845}1846}

1846 1847 

1847endpoint = Async::HTTP::Endpoint.parse("wss://api.openai.com/v1/responses", timeout: 10, alpn_protocols: ["http/1.1"])1848client = OpenAI::Client.new

1848headers = { "Authorization" => "Bearer #{ENV.fetch("OPENAI_API_KEY")}" }

1849Sync do |task|1849Sync do |task|

1850 task.with_timeout(120) do1850 task.with_timeout(120) do

1851 Async::WebSocket::Client.connect(endpoint, headers: headers) do |connection|1851 client.responses.connect(request_options: { timeout: 10 }) do |connection|

1852 connection.write(1852 connection.response.create(

1853 JSON.generate(1853 stream_id: "main", model: "gpt-6-astra", store: false,

1854 type: "response.create", stream_id: "main", model: "gpt-6-astra", store: false,

1855 input: [1854 input: [

1856 {1855 {

1857 role: "user",1856 role: "user",


1860 ],1859 ],

1861 tools: [test_log_tool, code_search_tool]1860 tools: [test_log_tool, code_search_tool]

1862 )1861 )

1863 )1862 puts(JSON.pretty_generate(wait_for_response(connection).output.map(&:to_h)))

1864 connection.flush

1865 puts(JSON.pretty_generate(wait_for_response(connection).fetch("output")))

1866 end1863 end

1867 end1864 end

1868end1865end

Details

269```269```

270 270 

271```ruby271```ruby

272require "csv"

272require "fileutils"273require "fileutils"

273require "json"274require "json"

274require "openai"275require "openai"


281 input: reviews.map { |review| review.tr("\n", " ") }282 input: reviews.map { |review| review.tr("\n", " ") }

282)283)

283 284 

284csv_field = ->(value) { %("#{value.gsub('"', '""')}") }

285rows = response.data.map.with_index do |embedding, index|

286 [csv_field.call(reviews.fetch(index)), csv_field.call(JSON.generate(embedding.embedding))].join(",")

287end

288 

289FileUtils.mkdir_p("output")285FileUtils.mkdir_p("output")

290File.write("output/embedded_1k_reviews.csv", (["combined,ada_embedding"] + rows).join("\n") + "\n")286CSV.open("output/embedded_1k_reviews.csv", "w") do |csv|

287 csv << ["combined", "ada_embedding"]

288 response.data.each do |embedding|

289 csv << [reviews.fetch(embedding.index), JSON.generate(embedding.embedding)]

290 end

291end

291```292```

292 293 

293 294 


903 904 

904```python905```python

905def recommendations_from_strings(906def recommendations_from_strings(

906 strings: List[str],907 strings: list[str],

907 index_of_source_string: int,908 index_of_source_string: int,

908 model="text-embedding-3-small",909 model="text-embedding-3-small",

909) -> List[int]:910) -> list[int]:

910 """Return nearest neighbors of a given string."""911 """Return nearest neighbors of a given string."""

911 912 

912 # get embeddings for all strings913 # get embeddings for all strings

Details

119const openai = new OpenAI();119const openai = new OpenAI();

120 120 

121// 1. Define a list of callable tools for the model121// 1. Define a list of callable tools for the model

122/** @type {OpenAI.Responses.Tool[]} */122 

123const tools = [123const tools = [

124 {124 {

125 type: "function",125 type: "function",


145}145}

146 146 

147// Create a running input list we will add to over time147// Create a running input list we will add to over time

148/** @type {OpenAI.Responses.ResponseInput} */148 

149let input = [149let input = [

150 { role: "user", content: "What is my horoscope? I am an Aquarius." },150 { role: "user", content: "What is my horoscope? I am an Aquarius." },

151];151];


1154 1154 

1155const openai = new OpenAI();1155const openai = new OpenAI();

1156 1156 

1157/** @type {OpenAI.Responses.Tool[]} */

1158const tools = [1157const tools = [

1159 {1158 {

1160 type: "function",1159 type: "function",

Details

1044from __future__ import annotations1044from __future__ import annotations

1045 1045 

1046import pathlib1046import pathlib

1047from collections.abc import Callable

1047from dataclasses import dataclass, field1048from dataclasses import dataclass, field

1048from enum import Enum1049from enum import Enum

1049from typing import (

1050 Callable,

1051 Dict,

1052 List,

1053 Optional,

1054 Tuple,

1055 Union,

1056)

1057 1050 

1058 1051 

1059# --------------------------------------------------------------------------- #1052# --------------------------------------------------------------------------- #


1068@dataclass1061@dataclass

1069class FileChange:1062class FileChange:

1070 type: ActionType1063 type: ActionType

1071 old_content: Optional[str] = None1064 old_content: str | None = None

1072 new_content: Optional[str] = None1065 new_content: str | None = None

1073 move_path: Optional[str] = None1066 move_path: str | None = None

1074 1067 

1075 1068 

1076@dataclass1069@dataclass

1077class Commit:1070class Commit:

1078 changes: Dict[str, FileChange] = field(default_factory=dict)1071 changes: dict[str, FileChange] = field(default_factory=dict)

1079 1072 

1080 1073 

1081# --------------------------------------------------------------------------- #1074# --------------------------------------------------------------------------- #


1091@dataclass1084@dataclass

1092class Chunk:1085class Chunk:

1093 orig_index: int = -11086 orig_index: int = -1

1094 del_lines: List[str] = field(default_factory=list)1087 del_lines: list[str] = field(default_factory=list)

1095 ins_lines: List[str] = field(default_factory=list)1088 ins_lines: list[str] = field(default_factory=list)

1096 1089 

1097 1090 

1098@dataclass1091@dataclass

1099class PatchAction:1092class PatchAction:

1100 type: ActionType1093 type: ActionType

1101 new_file: Optional[str] = None1094 new_file: str | None = None

1102 chunks: List[Chunk] = field(default_factory=list)1095 chunks: list[Chunk] = field(default_factory=list)

1103 move_path: Optional[str] = None1096 move_path: str | None = None

1104 1097 

1105 1098 

1106@dataclass1099@dataclass

1107class Patch:1100class Patch:

1108 actions: Dict[str, PatchAction] = field(default_factory=dict)1101 actions: dict[str, PatchAction] = field(default_factory=dict)

1109 1102 

1110 1103 

1111# --------------------------------------------------------------------------- #1104# --------------------------------------------------------------------------- #


1113# --------------------------------------------------------------------------- #1106# --------------------------------------------------------------------------- #

1114@dataclass1107@dataclass

1115class Parser:1108class Parser:

1116 current_files: Dict[str, str]1109 current_files: dict[str, str]

1117 lines: List[str]1110 lines: list[str]

1118 index: int = 01111 index: int = 0

1119 patch: Patch = field(default_factory=Patch)1112 patch: Patch = field(default_factory=Patch)

1120 fuzz: int = 01113 fuzz: int = 0


1131 return line.rstrip("\r")1124 return line.rstrip("\r")

1132 1125 

1133 # ------------- scanning convenience ----------------------------------- #1126 # ------------- scanning convenience ----------------------------------- #

1134 def is_done(self, prefixes: Optional[Tuple[str, ...]] = None) -> bool:1127 def is_done(self, prefixes: tuple[str, ...] | None = None) -> bool:

1135 if self.index >= len(self.lines):1128 if self.index >= len(self.lines):

1136 return True1129 return True

1137 if (1130 if (


1142 return True1135 return True

1143 return False1136 return False

1144 1137 

1145 def startswith(self, prefix: Union[str, Tuple[str, ...]]) -> bool:1138 def startswith(self, prefix: str | tuple[str, ...]) -> bool:

1146 return self._norm(self._cur_line()).startswith(prefix)1139 return self._norm(self._cur_line()).startswith(prefix)

1147 1140 

1148 def read_str(self, prefix: str) -> str:1141 def read_str(self, prefix: str) -> str:


1263 return action1256 return action

1264 1257 

1265 def _parse_add_file(self) -> PatchAction:1258 def _parse_add_file(self) -> PatchAction:

1266 lines: List[str] = []1259 lines: list[str] = []

1267 while not self.is_done(1260 while not self.is_done(

1268 ("*** End Patch", "*** Update File:", "*** Delete File:", "*** Add File:")1261 ("*** End Patch", "*** Update File:", "*** Delete File:", "*** Add File:")

1269 ):1262 ):


1278# Helper functions1271# Helper functions

1279# --------------------------------------------------------------------------- #1272# --------------------------------------------------------------------------- #

1280def find_context_core(1273def find_context_core(

1281 lines: List[str], context: List[str], start: int1274 lines: list[str], context: list[str], start: int

1282) -> Tuple[int, int]:1275) -> tuple[int, int]:

1283 if not context:1276 if not context:

1284 return start, 01277 return start, 0

1285 1278 


1300 1293 

1301 1294 

1302def find_context(1295def find_context(

1303 lines: List[str], context: List[str], start: int, eof: bool1296 lines: list[str], context: list[str], start: int, eof: bool

1304) -> Tuple[int, int]:1297) -> tuple[int, int]:

1305 if eof:1298 if eof:

1306 new_index, fuzz = find_context_core(lines, context, len(lines) - len(context))1299 new_index, fuzz = find_context_core(lines, context, len(lines) - len(context))

1307 if new_index != -1:1300 if new_index != -1:


1312 1305 

1313 1306 

1314def peek_next_section(1307def peek_next_section(

1315 lines: List[str], index: int1308 lines: list[str], index: int

1316) -> Tuple[List[str], List[Chunk], int, bool]:1309) -> tuple[list[str], list[Chunk], int, bool]:

1317 old: List[str] = []1310 old: list[str] = []

1318 del_lines: List[str] = []1311 del_lines: list[str] = []

1319 ins_lines: List[str] = []1312 ins_lines: list[str] = []

1320 chunks: List[Chunk] = []1313 chunks: list[Chunk] = []

1321 mode = "keep"1314 mode = "keep"

1322 orig_index = index1315 orig_index = index

1323 1316 


1397 if action.type is not ActionType.UPDATE:1390 if action.type is not ActionType.UPDATE:

1398 raise DiffError("_get_updated_file called with non-update action")1391 raise DiffError("_get_updated_file called with non-update action")

1399 orig_lines = text.split("\n")1392 orig_lines = text.split("\n")

1400 dest_lines: List[str] = []1393 dest_lines: list[str] = []

1401 orig_index = 01394 orig_index = 0

1402 1395 

1403 for chunk in action.chunks:1396 for chunk in action.chunks:


1420 return "\n".join(dest_lines)1413 return "\n".join(dest_lines)

1421 1414 

1422 1415 

1423def patch_to_commit(patch: Patch, orig: Dict[str, str]) -> Commit:1416def patch_to_commit(patch: Patch, orig: dict[str, str]) -> Commit:

1424 commit = Commit()1417 commit = Commit()

1425 for path, action in patch.actions.items():1418 for path, action in patch.actions.items():

1426 if action.type is ActionType.DELETE:1419 if action.type is ActionType.DELETE:


1447# --------------------------------------------------------------------------- #1440# --------------------------------------------------------------------------- #

1448# User-facing helpers1441# User-facing helpers

1449# --------------------------------------------------------------------------- #1442# --------------------------------------------------------------------------- #

1450def text_to_patch(text: str, orig: Dict[str, str]) -> Tuple[Patch, int]:1443def text_to_patch(text: str, orig: dict[str, str]) -> tuple[Patch, int]:

1451 lines = text.splitlines() # preserves blank lines, no strip()1444 lines = text.splitlines() # preserves blank lines, no strip()

1452 if (1445 if (

1453 len(lines) < 21446 len(lines) < 2


1461 return parser.patch, parser.fuzz1454 return parser.patch, parser.fuzz

1462 1455 

1463 1456 

1464def identify_files_needed(text: str) -> List[str]:1457def identify_files_needed(text: str) -> list[str]:

1465 lines = text.splitlines()1458 lines = text.splitlines()

1466 return [1459 return [

1467 line[len("*** Update File: ") :]1460 line[len("*** Update File: ") :]


1474 ]1467 ]

1475 1468 

1476 1469 

1477def identify_files_added(text: str) -> List[str]:1470def identify_files_added(text: str) -> list[str]:

1478 lines = text.splitlines()1471 lines = text.splitlines()

1479 return [1472 return [

1480 line[len("*** Add File: ") :]1473 line[len("*** Add File: ") :]


1486# --------------------------------------------------------------------------- #1479# --------------------------------------------------------------------------- #

1487# File-system helpers1480# File-system helpers

1488# --------------------------------------------------------------------------- #1481# --------------------------------------------------------------------------- #

1489def load_files(paths: List[str], open_fn: Callable[[str], str]) -> Dict[str, str]:1482def load_files(paths: list[str], open_fn: Callable[[str], str]) -> dict[str, str]:

1490 return {path: open_fn(path) for path in paths}1483 return {path: open_fn(path) for path in paths}

1491 1484 

1492 1485 

Details

73- **Reasoning effort:** Use `reasoning.effort` to choose between `low`, `medium`, `high`, or `xhigh`. The default is `medium`, but many workloads will perform well with `low`. Reserve `none` for use cases where low latency is more important than intelligence. See [Reasoning Models](https://developers.openai.com/api/docs/guides/reasoning) for detailed recommendations.73- **Reasoning effort:** Use `reasoning.effort` to choose between `low`, `medium`, `high`, or `xhigh`. The default is `medium`, but many workloads will perform well with `low`. Reserve `none` for use cases where low latency is more important than intelligence. See [Reasoning Models](https://developers.openai.com/api/docs/guides/reasoning) for detailed recommendations.

74- **Verbosity:** Use `text.verbosity` to control output length. Treat final answer length as separate from reasoning quality; specify word budgets, section counts, table widths, or JSON-only output where needed.74- **Verbosity:** Use `text.verbosity` to control output length. Treat final answer length as separate from reasoning quality; specify word budgets, section counts, table widths, or JSON-only output where needed.

75- **Structured Outputs:** Avoid describing the expected output schema in the prompt. Use [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) for automatic validation and increased accuracy.75- **Structured Outputs:** Avoid describing the expected output schema in the prompt. Use [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) for automatic validation and increased accuracy.

76- **Prompt caching:** [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) works automatically for eligible long prompts and can reduce latency and input-token cost. To maximize cache hits, keep stable content at the beginning of the request. Put dynamic user-specific context near the end. Track `usage.prompt_tokens_details.cached_tokens` to measure reuse. Use an optional [`prompt_cache_key`](https://developers.openai.com/api/docs/guides/prompt-caching#separate-prompts-with-cache-keys) to maintain separate cache accounting for customers or users, making cached token usage and billing easier to explain for each group. This also helps prevent cache-hit probing across users.76- **Prompt caching:** [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) works automatically for eligible long prompts and can reduce latency and input-token cost. To maximize cache hits, keep stable content at the beginning of the request. Put dynamic user-specific context near the end. Track `usage.prompt_tokens_details.cached_tokens` to measure reuse. Use a stable [`prompt_cache_key`](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-keys) for requests that share a reusable prefix. The key helps route related requests to the same cache and is important for optimizing cache hit rates on GPT-5.5. For busy groups, follow the [guidance for distributing traffic across more keys](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-keys).

77- **Tool calling:** GPT-5.5 supports the same tool-calling patterns as GPT-5.4, including function tools and tool-heavy agent workflows. Put most tool-specific guidance in the tool descriptions themselves: what the tool does, when to use it, required inputs, side effects, retry safety, and common error modes. Add tool-specific context to system instructions only when it applies across tools or materially changes the agent's operating policy.77- **Tool calling:** GPT-5.5 supports the same tool-calling patterns as GPT-5.4, including function tools and tool-heavy agent workflows. Put most tool-specific guidance in the tool descriptions themselves: what the tool does, when to use it, required inputs, side effects, retry safety, and common error modes. Add tool-specific context to system instructions only when it applies across tools or materially changes the agent's operating policy.

78- **Hosted tools and tool search:** Prefer [OpenAI-hosted tools](https://developers.openai.com/api/docs/guides/tools) where they fit the workflow, such as web search, file search, code interpreter, image generation, and computer use. Hosted tools reduce custom orchestration burden and keep common tool patterns aligned with the Responses API and Agents SDK. Use custom function tools when you need to call your own systems, enforce domain-specific side effects, or expose internal business workflows. For large tool catalogs, consider using [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search) to defer tool definitions and load only the relevant subset.78- **Hosted tools and tool search:** Prefer [OpenAI-hosted tools](https://developers.openai.com/api/docs/guides/tools) where they fit the workflow, such as web search, file search, code interpreter, image generation, and computer use. Hosted tools reduce custom orchestration burden and keep common tool patterns aligned with the Responses API and Agents SDK. Use custom function tools when you need to call your own systems, enforce domain-specific side effects, or expose internal business workflows. For large tool catalogs, consider using [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search) to defer tool definitions and load only the relevant subset.

79- **Tool preambles:** Preambles can improve chat UX because the user sees an initial, useful status update before the model generates the final response. They also make tool use easier to follow: the model can state what it's about to check or do, then continue from that same assistant state after tool results arrive.79- **Tool preambles:** Preambles can improve chat UX because the user sees an initial, useful status update before the model generates the final response. They also make tool use easier to follow: the model can state what it's about to check or do, then continue from that same assistant state after tool results arrive.

Details

69Include prior text messages in `session.input` when you create the session. For example, add this `input` field to your [session creation configuration](https://developers.openai.com/api/docs/guides/live#connect-your-first-session):69Include prior text messages in `session.input` when you create the session. For example, add this `input` field to your [session creation configuration](https://developers.openai.com/api/docs/guides/live#connect-your-first-session):

70 70 

71```javascript71```javascript

72/** @type {import("openai/resources/live/live").SessionConfig} */

73```72```

74 73 

75```python74```python


118Each event takes plain-string `content` of up to 500 tokens and a required `delegation_id`. Use `null` for session-wide context. For example, send this after your application has verified the user's acceptance and started the lookup:117Each event takes plain-string `content` of up to 500 tokens and a required `delegation_id`. Use `null` for session-wide context. For example, send this after your application has verified the user's acceptance and started the lookup:

119 118 

120```javascript119```javascript

121/**

122 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

123 */

124export function sendUpdate(connection) {120export function sendUpdate(connection) {

125 connection.send({121 connection.send({

126 type: "session.thinking.append",122 type: "session.thinking.append",


224Send `session.input_audio.mute` to mute input without ending the session:220Send `session.input_audio.mute` to mute input without ending the session:

225 221 

226```javascript222```javascript

227/**

228 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

229 */

230export function sendUpdate(connection) {223export function sendUpdate(connection) {

231 connection.send({224 connection.send({

232 type: "session.input_audio.mute",225 type: "session.input_audio.mute",


272Use `session.instructions.append` to request specific spoken wording for a disclosure. `session.commentary.append` may paraphrase the text. After `session.started`, for example, send:265Use `session.instructions.append` to request specific spoken wording for a disclosure. `session.commentary.append` may paraphrase the text. After `session.started`, for example, send:

273 266 

274```javascript267```javascript

275/**

276 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

277 */

278export function sendUpdate(connection) {268export function sendUpdate(connection) {

279 connection.send({269 connection.send({

280 type: "session.instructions.append",270 type: "session.instructions.append",

Details

48Add this delegation configuration when [creating your Live session](https://developers.openai.com/api/docs/guides/live). Choose the Responses model independently of the voice model:48Add this delegation configuration when [creating your Live session](https://developers.openai.com/api/docs/guides/live). Choose the Responses model independently of the voice model:

49 49 

50```javascript50```javascript

51/** @type {import("openai/resources/live/live").SessionConfig} */

52```51```

53 52 

54```python53```python


117After executing the authorized operation, append the result as a Responses item:116After executing the authorized operation, append the result as a Responses item:

118 117 

119```javascript118```javascript

120/**

121 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

122 */

123export function sendUpdate(connection) {119export function sendUpdate(connection) {

124 connection.send({120 connection.send({

125 type: "response.item.create",121 type: "response.item.create",


157Then explicitly continue the response:153Then explicitly continue the response:

158 154 

159```javascript155```javascript

160/**

161 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

162 */

163export function sendUpdate(connection) {156export function sendUpdate(connection) {

164 connection.send({157 connection.send({

165 type: "response.create",158 type: "response.create",


196Set `delegation` when [creating your Live session](https://developers.openai.com/api/docs/guides/live):189Set `delegation` when [creating your Live session](https://developers.openai.com/api/docs/guides/live):

197 190 

198```javascript191```javascript

199/** @type {import("openai/resources/live/live").SessionConfig} */

200```192```

201 193 

202```python194```python


242Return a result using that ID:234Return a result using that ID:

243 235 

244```javascript236```javascript

245/**

246 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

247 */

248export function sendUpdate(connection) {237export function sendUpdate(connection) {

249 connection.send({238 connection.send({

250 type: "session.commentary.append",239 type: "session.commentary.append",


320For quiet progress during a client-managed task:309For quiet progress during a client-managed task:

321 310 

322```javascript311```javascript

323/**

324 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

325 */

326export function sendUpdate(connection) {312export function sendUpdate(connection) {

327 connection.send({313 connection.send({

328 type: "session.thinking.append",314 type: "session.thinking.append",


352For a confirmed booking, send the result the user should hear:338For a confirmed booking, send the result the user should hear:

353 339 

354```javascript340```javascript

355/**

356 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

357 */

358export function sendUpdate(connection) {341export function sendUpdate(connection) {

359 connection.send({342 connection.send({

360 type: "session.commentary.append",343 type: "session.commentary.append",


386For example, after your application blocks a request under its guardrails, you can redirect the conversation:369For example, after your application blocks a request under its guardrails, you can redirect the conversation:

387 370 

388```javascript371```javascript

389/**

390 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

391 */

392export function sendUpdate(connection) {372export function sendUpdate(connection) {

393 connection.send({373 connection.send({

394 type: "session.instructions.append",374 type: "session.instructions.append",


465With Responses delegation, queue a user message for the backend:445With Responses delegation, queue a user message for the backend:

466 446 

467```javascript447```javascript

468/**

469 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

470 */

471export function sendUpdate(connection) {448export function sendUpdate(connection) {

472 connection.send({449 connection.send({

473 type: "response.item.create",450 type: "response.item.create",

Details

152**After: GPT-Live result**152**After: GPT-Live result**

153 153 

154```javascript154```javascript

155/**

156 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

157 */

158export function sendUpdate(connection) {155export function sendUpdate(connection) {

159 connection.send({156 connection.send({

160 type: "response.item.create",157 type: "response.item.create",


192After submitting every required function result, continue the backend:189After submitting every required function result, continue the backend:

193 190 

194```javascript191```javascript

195/**

196 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

197 */

198export function sendUpdate(connection) {192export function sendUpdate(connection) {

199 connection.send({193 connection.send({

200 type: "response.create",194 type: "response.create",


264Connect a client delegation to your agent258Connect a client delegation to your agent

265 259 

266```javascript260```javascript

267/**

268 * @typedef {{revision: number, recentConversation: string, task: string}} Context

269 * @typedef {import("openai/resources/live/live").ServerEvent} Notice

270 * @typedef {import("openai/resources/live/live").CommentaryAppendEvent} Update

271 * @param {Notice} event

272 * @param {{readContext: () => Context | null, runAgent: (context: Context) => Promise<string>, currentRevision: () => number, send: (event: Update) => void}} app

273 */

274async function handleDelegation(event, app) {261async function handleDelegation(event, app) {

275 if (262 if (

276 event.type !== "session.delegation.created" ||263 event.type !== "session.delegation.created" ||

Details

219Reuse simple message input219Reuse simple message input

220 220 

221```javascript221```javascript

222/** @type {OpenAI.ChatCompletionMessageParam[] & OpenAI.Responses.ResponseInput} */

223const context = [222const context = [

224 { role: "system", content: "You are a helpful assistant." },223 { role: "system", content: "You are a helpful assistant." },

225 { role: "user", content: "Hello!" },224 { role: "user", content: "Hello!" },


711 Multi-turn conversation710 Multi-turn conversation

712 711 

713```javascript712```javascript

714/** @type {OpenAI.ChatCompletionMessageParam[]} */

715let messages = [713let messages = [

716 { role: "system", content: "You are a helpful assistant." },714 { role: "system", content: "You are a helpful assistant." },

717 { role: "user", content: "What is the capital of France?" },715 { role: "user", content: "What is the capital of France?" },


873```javascript871```javascript

874import { toResponseInputItems } from "openai/lib/responses/ResponseInputItems";872import { toResponseInputItems } from "openai/lib/responses/ResponseInputItems";

875 873 

876/** @type {OpenAI.Responses.ResponseInput} */

877let context = [{ role: "user", content: "What is the capital of France?" }];874let context = [{ role: "user", content: "What is the capital of France?" }];

878 875 

879const res1 = await client.responses.create({876const res1 = await client.responses.create({

Details

201 201 

202- Current machine load and available capacity.202- Current machine load and available capacity.

203- A hash of the initial tokens after the hidden OpenAI content, including tool definitions when present. The number of tokens hashed varies by model.203- A hash of the initial tokens after the hidden OpenAI content, including tool definitions when present. The number of tokens hashed varies by model.

204- An optional [`prompt_cache_key`](#prompt-cache-keys), which separates cache reuse between groups of requests.204- A supplied [`prompt_cache_key`](#prompt-cache-keys), which separates cache reuse between groups of requests and helps optimize cache routing on models before GPT-5.6.

205 205 

206 206 

207 207 


213 213 

214 214 

215 215 

216[`prompt_cache_key`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20prompt_cache_key%20%3E%20%28schema%29) is an optional control for maintaining separate cache accounting for customers or users within your application. OpenAI handles cache routing automatically; you can omit the key for normal caching.216On models before GPT-5.6, use a stable [`prompt_cache_key`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20prompt_cache_key%20%3E%20%28schema%29) for requests that share a reusable prefix to help route related requests to the same cache. For busy groups, aim for about 15 requests per minute in total across all prefixes using each key. Partition higher-volume traffic across multiple keys using a stable, deterministic mapping. Keep related requests on the same `prompt_cache_key` so they can reuse its cache. Keys influence routing; they do not pin requests to a machine or guarantee a cache hit.

217 

218On GPT-5.6 and later, OpenAI handles cache routing automatically; the key is not needed to optimize caching. You can use separate keys to maintain separate cache accounting for customers or users within your application.

217 219 

218Using separate keys can make cached token usage and billing easier to explain for each customer or user. For example, separate keys help prevent cache-hit probing across users: submitting candidate prompts and observing cache hits to learn whether matching content was previously cached. See [Separate cache accounting with keys](#separate-prompts-with-cache-keys).220Using separate keys can make cached token usage and billing easier to explain for each customer or user. For example, separate keys help prevent cache-hit probing across users: submitting candidate prompts and observing cache hits to learn whether matching content was previously cached. See [Separate cache accounting with keys](#separate-prompts-with-cache-keys).

219 221 


229| -------------------------- | --------------------------------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------- |231| -------------------------- | --------------------------------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------- |

230| Implicit breakpoints | At the end of the latest eligible message. | Spaced at regular 2,048-token intervals. | Spaced at regular, model-dependent intervals. |232| Implicit breakpoints | At the end of the latest eligible message. | Spaced at regular 2,048-token intervals. | Spaced at regular, model-dependent intervals. |

231| Explicit breakpoints | Supported | Not supported | Not supported |233| Explicit breakpoints | Supported | Not supported | Not supported |

234| `prompt_cache_key` | Optional for separate cache accounting | Use a stable key to optimize cache routing | Use a stable key to optimize cache routing |

232| Minimum cacheable prefix | 1,024 visible input tokens | Varies by request settings | Varies by request settings |235| Minimum cacheable prefix | 1,024 visible input tokens | Varies by request settings | Varies by request settings |

233| Cached-token reporting | Exact eligible boundary, excluding hidden tokens | Excludes hidden tokens and rounds down to a multiple of 128 | Excludes hidden tokens and rounds down to a multiple of 128 |236| Cached-token reporting | Exact eligible boundary, excluding hidden tokens | Excludes hidden tokens and rounds down to a multiple of 128 | Excludes hidden tokens and rounds down to a multiple of 128 |

234| Cache read charge | 0.1× the uncached input-token rate | Model-dependent cached-input rate | Model-dependent cached-input rate |237| Cache read charge | 0.1× the uncached input-token rate | Model-dependent cached-input rate | Model-dependent cached-input rate |


259 262 

260## How to optimize prompt caching263## How to optimize prompt caching

261 264 

262Focus on [preserving conversation history](#preserve-conversation-history), [keeping tool definitions stable](#manage-tools-with-append-only-updates), and choosing where caching occurs. Use [`prompt_cache_options.mode` and `prompt_cache_breakpoint`](#choose-a-caching-mode) to control cache breakpoints. If your application needs separate cache accounting for customers, you can also use an optional [`prompt_cache_key`](#separate-prompts-with-cache-keys).265Focus on [preserving conversation history](#preserve-conversation-history), [keeping tool definitions stable](#manage-tools-with-append-only-updates), and choosing where caching occurs. On GPT-5.6 and later, use [`prompt_cache_options.mode` and `prompt_cache_breakpoint`](#choose-a-caching-mode) to control cache breakpoints. You can also use an optional [`prompt_cache_key`](#separate-prompts-with-cache-keys) if your application needs separate cache accounting for customers. On models before GPT-5.6, use a stable `prompt_cache_key` to optimize cache routing for requests that share a reusable prefix.

263 266 

264 267 

265 268 


412 415 

413 416 

414 417 

415Use `prompt_cache_key` when you want to maintain separate cache accounting for customers, users, or workspaces within your application. This can make cached token usage and billing easier to explain within each group. The key is optional and is not needed to optimize caching.418On GPT-5.6 and later, use `prompt_cache_key` when you want to maintain separate cache accounting for customers, users, or workspaces within your application. This can make cached token usage and billing easier to explain within each group. The key is optional and is not needed to optimize caching on these models.

416 419 

417- **Choose how to separate cache accounting.** Assign a distinct key to each customer or user whose cache accounting should remain separate. For example, `support:customer_123` and `support:customer_456` maintain separate cache accounting for two customers, even when their requests contain the same prefix.420- **Choose how to separate cache accounting.** Assign a distinct key to each customer or user whose cache accounting should remain separate. For example, `support:customer_123` and `support:customer_456` maintain separate cache accounting for two customers, even when their requests contain the same prefix.

418- **Keep keys stable within each group.** Reuse the same key for a customer's related requests. Generate a separate key for a session or thread only when it needs its own cache accounting.421- **Keep keys stable within each group.** Reuse the same key for a customer's related requests. Generate a separate key for a session or thread only when it needs its own cache accounting.

419- **Apply keys consistently.** Use the customer's key across their requests to maintain separate cache accounting. This also helps prevent cache-hit probing across customers.422- **Apply keys consistently.** Use the customer's key across their requests to maintain separate cache accounting. This also helps prevent cache-hit probing across customers.

420 423 

424On models before GPT-5.6, `prompt_cache_key` is important for optimizing cache hit rates. Use a stable key for requests that share a reusable prefix to help route them to the same cache. For busy groups, follow the [guidance for distributing traffic across more keys](#prompt-cache-keys).

425 

421 426 

422 427 

423 428 


592 597 

593## Examples598## Examples

594 599 

600The following examples apply to GPT-5.6 and later models.

601 

595 602 

596 603 

597<a id="single-turn-llm-as-a-judge"></a>604<a id="single-turn-llm-as-a-judge"></a>

Details

349 349 

350const client = new OpenAI();350const client = new OpenAI();

351 351 

352/** @returns {OpenAI.Responses.ResponseInput} */

353function buildSupportPrompt({ customerName, issue }) {352function buildSupportPrompt({ customerName, issue }) {

354 return [353 return [

355 {354 {

Details

714 714 

715const client = new OpenAI();715const client = new OpenAI();

716 716 

717/** @type {OpenAI.Responses.ResponseInput} */

718const history = [717const history = [

719 {718 {

720 role: "user",719 role: "user",

Details

195 alpha: { estimated_weeks: 6, risk: "medium" },195 alpha: { estimated_weeks: 6, risk: "medium" },

196 beta: { estimated_weeks: 8, risk: "low" },196 beta: { estimated_weeks: 8, risk: "low" },

197};197};

198/** @type {import("openai/resources/beta/responses").BetaTool[]} */198 

199const tools = [199const tools = [

200 {200 {

201 type: "function",201 type: "function",


216 strict: true,216 strict: true,

217 },217 },

218];218];

219/**219 

220 * @type {Array<

221 * import("openai/resources/beta/responses").BetaResponseInputItem |

222 * import("openai/resources/beta/responses").BetaResponseOutputItem

223 * >}

224 */

225const history = [220const history = [

226 {221 {

227 role: "user",222 role: "user",


250 const stream = await client.beta.responses.create({245 const stream = await client.beta.responses.create({

251 model: "gpt-5.6-sol",246 model: "gpt-5.6-sol",

252 // Beta output items can be replayed as input on the next request.247 // Beta output items can be replayed as input on the next request.

253 input:248 input: history,

254 /** @type {import("openai/resources/beta/responses").BetaResponseInput} */ (

255 history

256 ),

257 tools,249 tools,

258 store: false,250 store: false,

259 multi_agent: {251 multi_agent: {

Details

126 126 

127client = OpenAI::Client.new127client = OpenAI::Client.new

128store = client.vector_stores.create(name: "Support FAQ")128store = client.vector_stores.create(name: "Support FAQ")

129source = Pathname("customer_policies.txt")129file = client.vector_stores.files.upload_and_poll(

130uploaded = client.files.create(file: source, purpose: :assistants)130 store.id,

131file = client.vector_stores.files.create(store.id, file_id: uploaded.id)131 file: Pathname("customer_policies.txt"),

132until [:completed, :failed, :cancelled].include?(file.status)132 timeout: 600

133 sleep(1)133)

134 file = client.vector_stores.files.retrieve(file.id, vector_store_id: store.id)134raise "File ingestion ended with status: #{file.status}" unless file.status == OpenAI::VectorStores::VectorStoreFile::Status::COMPLETED

135end

136 135 

137puts(store.id)136puts(store.id)

138```137```


1020require "pathname"1019require "pathname"

1021 1020 

1022client = OpenAI::Client.new1021client = OpenAI::Client.new

1023file = Pathname("customer_policies.txt")1022vector_store_file = client.vector_stores.files.upload_and_poll(

1024uploaded = client.files.create(file: file, purpose: :assistants)

1025vector_store_file = client.vector_stores.files.create(

1026 "vs_123",1023 "vs_123",

1027 file_id: uploaded.id1024 file: Pathname("customer_policies.txt"),

1025 timeout: 600

1028)1026)

1029until [:completed, :failed, :cancelled].include?(vector_store_file.status)1027raise "File ingestion ended with status: #{vector_store_file.status}" unless vector_store_file.status == OpenAI::VectorStores::VectorStoreFile::Status::COMPLETED

1030 sleep(1)1028 

1031 vector_store_file = client.vector_stores.files.retrieve(

1032 vector_store_file.id,

1033 vector_store_id: "vs_123"

1034 )

1035end

1036puts(vector_store_file.id)1029puts(vector_store_file.id)

1037```1030```

1038 1031 


1455require "openai"1448require "openai"

1456 1449 

1457client = OpenAI::Client.new1450client = OpenAI::Client.new

1458batch = client.vector_stores.file_batches.create(1451batch = client.vector_stores.file_batches.create_and_poll(

1459 "vs_123",1452 "vs_123",

1460 files: [1453 files: [

1461 {1454 {


1466 file_id: "file_456",1459 file_id: "file_456",

1467 chunking_strategy: {1460 chunking_strategy: {

1468 type: :static,1461 type: :static,

1462 static: {

1469 max_chunk_size_tokens: 1_200,1463 max_chunk_size_tokens: 1_200,

1470 chunk_overlap_tokens: 2001464 chunk_overlap_tokens: 200

1471 }1465 }

1472 }1466 }

1473 ]1467 }

1468 ],

1469 timeout: 600

1474)1470)

1475until [:completed, :failed, :cancelled].include?(batch.status)1471raise "File ingestion ended with status: #{batch.status}" unless batch.status == OpenAI::VectorStores::VectorStoreFileBatch::Status::COMPLETED

1476 sleep(1)1472 

1477 batch = client.vector_stores.file_batches.retrieve(1473raise "File ingestion failed for #{batch.file_counts.failed} file(s)" if batch.file_counts.failed.positive?

1478 batch.id,1474 

1479 vector_store_id: "vs_123"1475# Live validation of per-file batches returned default chunking despite overrides.

1480 )1476file = client.vector_stores.files.retrieve("file_456", vector_store_id: "vs_123")

1477strategy = file.chunking_strategy

1478unless strategy.is_a?(OpenAI::StaticFileChunkingStrategyObject) &&

1479 strategy.static.max_chunk_size_tokens == 1_200 &&

1480 strategy.static.chunk_overlap_tokens == 200

1481 raise "Requested chunking was not applied to #{file.id}: #{strategy.to_json}"

1481end1482end

1483 

1482puts(batch.status)1484puts(batch.status)

1483```1485```

1484 1486 

Details

152# Note this file gets uploaded to the OpenAI API as a grader152# Note this file gets uploaded to the OpenAI API as a grader

153from ast_grep_py import SgRoot153from ast_grep_py import SgRoot

154from pydantic import BaseModel, Field # type: ignore154from pydantic import BaseModel, Field # type: ignore

155from typing import Any, List, Optional155from typing import Any

156import re156import re

157 157 

158SUPPORTED_LANGUAGES = ['typescript', 'javascript', 'ts', 'js']158SUPPORTED_LANGUAGES = ['typescript', 'javascript', 'ts', 'js']


173class ASTGrepPattern(BaseModel):173class ASTGrepPattern(BaseModel):

174 file_path_mask: str = Field(..., description="The file path pattern to match against")174 file_path_mask: str = Field(..., description="The file path pattern to match against")

175 pattern: str = Field(..., description="The main AST grep pattern to search for")175 pattern: str = Field(..., description="The main AST grep pattern to search for")

176 additional_greps: Optional[List[str]] = Field(176 additional_greps: list[str] | None = Field(

177 default=None,177 default=None,

178 description="Additional patterns that must also be present in the matched code"178 description="Additional patterns that must also be present in the matched code"

179 )179 )

180 180 

181def extract_code_blocks(llm_output: str) -> List[CodeBlock]:181def extract_code_blocks(llm_output: str) -> list[CodeBlock]:

182 # Regular expression to match code blocks with optional language and path182 # Regular expression to match code blocks with optional language and path

183 try:183 try:

184 pattern = r"```(\w+\s+)?([\w./-]+)?\n([\s\S]*?)\n```"184 pattern = r"```(\w+\s+)?([\w./-]+)?\n([\s\S]*?)\n```"


246 return code_blocks246 return code_blocks

247 247 

248 248 

249def calculate_ast_grep_score(code_blocks: List[CodeBlock], ast_greps: Any) -> float:249def calculate_ast_grep_score(code_blocks: list[CodeBlock], ast_greps: Any) -> float:

250 # Convert ast_greps to list if it's a dict250 # Convert ast_greps to list if it's a dict

251 if isinstance(ast_greps, dict):251 if isinstance(ast_greps, dict):

252 ast_greps = [ast_greps]252 ast_greps = [ast_greps]

253 253 

254 # Parse each grep pattern into the Pydantic model254 # Parse each grep pattern into the Pydantic model

255 parsed_patterns: List[ASTGrepPattern] = []255 parsed_patterns: list[ASTGrepPattern] = []

256 for grep in ast_greps:256 for grep in ast_greps:

257 try:257 try:

258 pattern = ASTGrepPattern(**grep)258 pattern = ASTGrepPattern(**grep)


400 code_start = output_text.find('<code>')400 code_start = output_text.find('<code>')

401 code_end = output_text.find('</code>')401 code_end = output_text.find('</code>')

402 code_to_grade: str = output_text[code_start + len('<code>'):code_end].strip()402 code_to_grade: str = output_text[code_start + len('<code>'):code_end].strip()

403 code_blocks: List[CodeBlock] = []403 code_blocks: list[CodeBlock] = []

404 try:404 try:

405 code_blocks = extract_code_blocks(code_to_grade)405 code_blocks = extract_code_blocks(code_to_grade)

406 except Exception as e:406 except Exception as e:

Details

317 317 

318const agentRef = fs.readFileSync("fixtures/agent.wav").toString("base64");318const agentRef = fs.readFileSync("fixtures/agent.wav").toString("base64");

319 319 

320const transcript = /** @type {OpenAI.Audio.TranscriptionDiarized} */ (320const transcript = await openai.audio.transcriptions.create({

321 await openai.audio.transcriptions.create({

322 file: fs.createReadStream("fixtures/meeting.wav"),321 file: fs.createReadStream("fixtures/meeting.wav"),

323 model: "gpt-4o-transcribe-diarize",322 model: "gpt-4o-transcribe-diarize",

324 response_format: "diarized_json",323 response_format: "diarized_json",

325 chunking_strategy: "auto",324 chunking_strategy: "auto",

326 known_speaker_names: ["agent"],325 known_speaker_names: ["agent"],

327 known_speaker_references: ["data:audio/wav;base64," + agentRef],326 known_speaker_references: ["data:audio/wav;base64," + agentRef],

328 })327});

329);

330 328 

331for (const segment of transcript.segments) {329for (const segment of transcript.segments) {

332 if (!("speaker" in segment)) continue;330 if (!("speaker" in segment)) continue;

guides/steering.md +22 −34

Details

192 192 

193```ruby193```ruby

194require "async"194require "async"

195require "async/http/endpoint"195require "openai"

196require "async/websocket/client"

197require "json"

198 196 

199endpoint = Async::HTTP::Endpoint.parse("wss://api.openai.com/v1/responses", timeout: 10, alpn_protocols: ["http/1.1"])197client = OpenAI::Client.new

200headers = { "Authorization" => "Bearer #{ENV.fetch("OPENAI_API_KEY")}" }

201Sync do |task|198Sync do |task|

202 task.with_timeout(120) do199 task.with_timeout(120) do

203 Async::WebSocket::Client.connect(endpoint, headers: headers) do |connection|200 client.responses.connect(request_options: { timeout: 10 }) do |connection|

204 connection.write(201 connection.response.create(

205 JSON.generate(202 model: "gpt-6-astra", reasoning: { effort: "medium" },

206 type: "response.create", model: "gpt-6-astra", reasoning: { effort: "medium" },

207 input: "Draft a project plan for building a task-tracking app."203 input: "Draft a project plan for building a task-tracking app."

208 )204 )

209 )

210 connection.flush

211 state = {}205 state = {}

212 while (message = connection.read)206 while (event = connection.receive)

213 event = JSON.parse(message.to_str)207 case event

214 response = event["response"]208 when OpenAI::Responses::ResponseCreatedEvent

215 case event.fetch("type")209 response = event.response

216 when "response.created"

217 if !state[:initial_id]210 if !state[:initial_id]

218 state[:initial_id] = response.fetch("id")211 state[:initial_id] = response.id

219 connection.write(212 connection.send_event(

220 JSON.generate(

221 type: "response.steer", previous_response_id: state[:initial_id],213 type: "response.steer", previous_response_id: state[:initial_id],

222 input: "Keep the scope small enough for one developer to finish in two weeks."214 input: "Keep the scope small enough for one developer to finish in two weeks."

223 )215 )

224 )

225 connection.flush

226 else216 else

227 state[:successor_id] = response.fetch("id")217 state[:successor_id] = response.id

228 end218 end

229 when "response.steer.failed", "response.failed", "error"219 when OpenAI::Responses::ResponseSteerFailedEvent, OpenAI::Responses::ResponseFailedEvent, OpenAI::Responses::ResponsesServerEvent::ResponseWsError

230 raise "Steering failed: #{JSON.generate(event)}"220 raise "Steering failed: #{event.to_json}"

231 when "response.incomplete"221 when OpenAI::Responses::ResponseIncompleteEvent

232 unless response.fetch("id") == state[:initial_id] && response.dig("incomplete_details", "reason") == "steered"222 response = event.response

233 raise "Response incomplete: #{JSON.generate(event)}"223 unless response.id == state[:initial_id] && response.incomplete_details&.reason.to_s == "steered"

224 raise "Response incomplete: #{event.to_json}"

234 end225 end

235 when "response.completed"226 when OpenAI::Responses::ResponseCompletedEvent

236 next unless state[:successor_id] && response.fetch("id") == state[:successor_id]227 response = event.response

237 228 next unless state[:successor_id] && response.id == state[:successor_id]

238 response.fetch("output").each do |item|

239 next unless item["type"] == "message"

240 229 

241 item.fetch("content").each { |part| puts(part.fetch("text")) if part["type"] == "output_text" }230 puts(response.output_text)

242 end

243 state[:completed] = true231 state[:completed] = true

244 break232 break

245 end233 end

Details

121. **Explicit refusals:** Safety-based model refusals are now programmatically detectable121. **Explicit refusals:** Safety-based model refusals are now programmatically detectable

131. **Simpler prompting:** No need for strongly worded prompts to achieve consistent formatting131. **Simpler prompting:** No need for strongly worded prompts to achieve consistent formatting

14 14 

15In addition to supporting JSON Schema in the REST API, the OpenAI SDKs for [Python](https://github.com/openai/openai-python/blob/main/helpers.md#structured-outputs-parsing-helpers) and [JavaScript](https://github.com/openai/openai-node/blob/master/helpers.md#structured-outputs-parsing-helpers) also make it easy to define object schemas using [Pydantic](https://docs.pydantic.dev/latest/) and [Zod](https://zod.dev/) respectively. Below, you can see how to extract information from unstructured text that conforms to a schema defined in code.15In addition to supporting JSON Schema in the REST API, the OpenAI libraries for [Python](https://github.com/openai/openai-python/blob/main/helpers.md#structured-outputs-parsing-helpers) and [JavaScript](https://github.com/openai/openai-node/blob/master/helpers.md#structured-outputs-parsing-helpers) also let you define object schemas using [`pydantic.BaseModel`](https://docs.pydantic.dev/latest/) and [`z.object`](https://zod.dev/) respectively. Below, you can see how to extract information from unstructured text that conforms to a schema defined in code.

16 

17The Ruby SDK supports schemas defined with Sorbet `T::Struct` and returns typed parsed results.

16 18 

17 19 

18 20 


236```238```

237 239 

238```ruby240```ruby

241# gem install openai sorbet-runtime

239require "openai"242require "openai"

243require "openai/helpers/sorbet"

244 

245class CalendarEvent < T::Struct

246 const :name, String

247 const :date, String

248 const :participants, T::Array[String]

249end

240 250 

241client = OpenAI::Client.new251client = OpenAI::Client.new

242event_schema = {252schema = OpenAI::StructuredOutput.from_sorbet(CalendarEvent)

243 type: :object,

244 properties: {

245 name: { type: :string },

246 date: { type: :string },

247 participants: {

248 type: :array,

249 items: { type: :string }

250 }

251 },

252 required: %w[name date participants],

253 additionalProperties: false

254}

255 253 

256response = client.responses.create(254response = client.responses.create(

257 model: "gpt-6-astra",255 model: "gpt-6-astra",


265 content: "Alice and Bob are going to a science fair on Friday."263 content: "Alice and Bob are going to a science fair on Friday."

266 }264 }

267 ],265 ],

268 text: {266 text: schema

269 format: {

270 type: :json_schema,

271 name: "event",

272 strict: true,

273 schema: event_schema

274 }

275 }

276)267)

277 268 

278puts(response.output_text)269raise "Response ended with status: #{response.status}" unless response.status == OpenAI::Responses::ResponseStatus::COMPLETED

270 

271message = response.output.grep(OpenAI::Responses::ResponseOutputMessage).fetch(0)

272output_text = message.content.grep(OpenAI::Responses::ResponseOutputText).first

273raise "No structured output returned (the model may have refused)" unless output_text

274 

275event = T.cast(output_text.parsed, CalendarEvent)

276puts(event.name, event.date, event.participants.join(", "))

279```277```

280 278 

281 279 


309 307 

310For example, if you are building a math tutoring application, you might want the assistant to respond to your user using a specific JSON Schema so that you can generate a UI that displays different parts of the model's output in distinct ways.308For example, if you are building a math tutoring application, you might want the assistant to respond to your user using a specific JSON Schema so that you can generate a UI that displays different parts of the model's output in distinct ways.

311 309 

312Put simply:310In practice:

313 311 

314 312 

315 313 


1195 1193 

1196```python1194```python

1197from enum import Enum1195from enum import Enum

1198from typing import List

1199 1196 

1200from openai import OpenAI1197from openai import OpenAI

1201from pydantic import BaseModel1198from pydantic import BaseModel


1220class UI(BaseModel):1217class UI(BaseModel):

1221 type: UIType1218 type: UIType

1222 label: str1219 label: str

1223 children: List["UI"]1220 children: list["UI"]

1224 attributes: List[Attribute]1221 attributes: list[Attribute]

1225 1222 

1226 1223 

1227UI.model_rebuild() # This is required to enable recursive types1224UI.model_rebuild() # This is required to enable recursive types


1706 1703 

1707```python1704```python

1708from enum import Enum1705from enum import Enum

1709from typing import Optional

1710 1706 

1711from openai import OpenAI1707from openai import OpenAI

1712from pydantic import BaseModel1708from pydantic import BaseModel


1722 1718 

1723class ContentCompliance(BaseModel):1719class ContentCompliance(BaseModel):

1724 is_violating: bool1720 is_violating: bool

1725 category: Optional[Category]1721 category: Category | None

1726 explanation_if_violating: Optional[str]1722 explanation_if_violating: str | None

1727 1723 

1728 1724 

1729response = client.responses.parse(1725response = client.responses.parse(


3331 3327 

3332#### Avoid JSON schema divergence3328#### Avoid JSON schema divergence

3333 3329 

3334To prevent your JSON Schema and corresponding types in your programming language from diverging, we strongly recommend using the native Pydantic/zod sdk support.3330To prevent your JSON Schema and corresponding types in your programming language from diverging, we strongly recommend using native SDK schema helpers where available.

3335 3331 

3336If you prefer to specify the JSON schema directly, you could add CI rules that flag when either the JSON schema or underlying data objects are edited, or add a CI step that auto-generates the JSON Schema from type definitions (or vice-versa).3332If you prefer to specify the JSON schema directly, you could add CI rules that flag when either the JSON schema or underlying data objects are edited, or add a CI step that automatically generates the JSON Schema from type definitions (or vice-versa).

3337 3333 

3338## Streaming3334## Streaming

3339 3335 


3390```3386```

3391 3387 

3392```python3388```python

3393from typing import List

3394 

3395from openai import OpenAI3389from openai import OpenAI

3396from pydantic import BaseModel3390from pydantic import BaseModel

3397 3391 

3398 3392 

3399class EntitiesModel(BaseModel):3393class EntitiesModel(BaseModel):

3400 attributes: List[str]3394 attributes: list[str]

3401 colors: List[str]3395 colors: list[str]

3402 animals: List[str]3396 animals: list[str]

3403 3397 

3404 3398 

3405client = OpenAI()3399client = OpenAI()

guides/tools.md +0 −2

Details

280 280 

281const client = new OpenAI();281const client = new OpenAI();

282 282 

283/** @type {OpenAI.Responses.NamespaceTool} */

284const crmNamespace = {283const crmNamespace = {

285 type: "namespace",284 type: "namespace",

286 name: "crm",285 name: "crm",


547import OpenAI from "openai";546import OpenAI from "openai";

548const client = new OpenAI();547const client = new OpenAI();

549 548 

550/** @type {OpenAI.Responses.Tool[]} */

551const tools = [549const tools = [

552 {550 {

553 type: "function",551 type: "function",

Details

183Apply the patch and return results183Apply the patch and return results

184 184 

185```javascript185```javascript

186/** @type {import("openai/resources/responses/responses").ResponseInput} */

187const results = patchCalls.map((call) => {186const results = patchCalls.map((call) => {

188 const { success, output } = applyOperation(call.operation);187 const { success, output } = applyOperation(call.operation);

189 188 


364import { applyDiff, Agent, run, applyPatchTool } from "@openai/agents";363import { applyDiff, Agent, run, applyPatchTool } from "@openai/agents";

365 364 

366class WorkspaceEditor {365class WorkspaceEditor {

367 /** @returns {Promise<import("@openai/agents").ApplyPatchResult>} */

368 async createFile(operation) {366 async createFile(operation) {

369 // convert the diff to the file content367 // convert the diff to the file content

370 const content = applyDiff("", operation.diff, "create");368 const content = applyDiff("", operation.diff, "create");


372 return { status: "completed", output: `Created ${operation.path}` };370 return { status: "completed", output: `Created ${operation.path}` };

373 }371 }

374 372 

375 /** @returns {Promise<import("@openai/agents").ApplyPatchResult>} */

376 async updateFile(operation) {373 async updateFile(operation) {

377 // read the file content from the file system374 // read the file content from the file system

378 const current = "";375 const current = "";


382 return { status: "completed", output: `Updated ${operation.path}` };379 return { status: "completed", output: `Updated ${operation.path}` };

383 }380 }

384 381 

385 /** @returns {Promise<import("@openai/agents").ApplyPatchResult>} */

386 async deleteFile(operation) {382 async deleteFile(operation) {

387 // delete the file from the file system383 // delete the file from the file system

388 return { status: "completed", output: `Deleted ${operation.path}` };384 return { status: "completed", output: `Deleted ${operation.path}` };

Details

144async function runComputerUse(endpoint, prompt, model = "gpt-6-astra") {144async function runComputerUse(endpoint, prompt, model = "gpt-6-astra") {

145 const client = new OpenAI();145 const client = new OpenAI();

146 const sessionId = randomUUID();146 const sessionId = randomUUID();

147 /** @type {OpenAI.Responses.Tool[]} */147 

148 const tools = [148 const tools = [

149 {149 {

150 type: "function",150 type: "function",


164 strict: true,164 strict: true,

165 },165 },

166 ];166 ];

167 /** @type {OpenAI.Responses.ResponseInput} */167 

168 let nextInput = [{ role: "user", content: prompt }];168 let nextInput = [{ role: "user", content: prompt }];

169 let previousResponseId;169 let previousResponseId;

170 170 


465const client = new OpenAI();465const client = new OpenAI();

466 466 

467async function sendComputerScreenshot(response, callId, screenshotBase64) {467async function sendComputerScreenshot(response, callId, screenshotBase64) {

468 const output = /** @type {const} */ ({468 const output = {

469 type: "computer_screenshot",469 type: "computer_screenshot",

470 image_url: `data:image/png;base64,${screenshotBase64}`,470 image_url: `data:image/png;base64,${screenshotBase64}`,

471 detail: "original",471 detail: "original",

472 });472 };

473 473 

474 return await client.responses.create({474 return await client.responses.create({

475 model: "gpt-5.6-sol",475 model: "gpt-5.6-sol",

Details

1456 1456 

1457 const screenshot = await captureScreenshot(target);1457 const screenshot = await captureScreenshot(target);

1458 const screenshotBase64 = Buffer.from(screenshot).toString("base64");1458 const screenshotBase64 = Buffer.from(screenshot).toString("base64");

1459 const output = /** @type {const} */ ({1459 const output = {

1460 type: "computer_screenshot",1460 type: "computer_screenshot",

1461 image_url: `data:image/png;base64,${screenshotBase64}`,1461 image_url: `data:image/png;base64,${screenshotBase64}`,

1462 detail: "original",1462 detail: "original",

1463 });1463 };

1464 1464 

1465 response = await client.responses.create({1465 response = await client.responses.create({

1466 model: "gpt-5.6-sol",1466 model: "gpt-5.6-sol",


1909 )1909 )

1910 .nonempty();1910 .nonempty();

1911 1911 

1912/** @returns {Promise<import("openai/resources/responses/responses").ResponseFunctionCallOutputItemList>} */

1913async function executeInSandbox(code, sessionId, endpoint) {1912async function executeInSandbox(code, sessionId, endpoint) {

1914 console.log(code);1913 console.log(code);

1915 const terminal = readline.createInterface({1914 const terminal = readline.createInterface({

Details

234 get_demand: async ({ sku }) => ({ sku, requested_units: 31 }),234 get_demand: async ({ sku }) => ({ sku, requested_units: 31 }),

235};235};

236 236 

237/** @type {OpenAI.Responses.Tool[]} */

238const tools = [237const tools = [

239 {238 {

240 type: "function",239 type: "function",


285 { type: "programmatic_tool_calling" },284 { type: "programmatic_tool_calling" },

286];285];

287 286 

288/** @type {OpenAI.Responses.ResponseInput} */

289const input = [287const input = [

290 {288 {

291 role: "user",289 role: "user",


326 if (!run) throw new Error(`Unknown tool: ${call.name}`);324 if (!run) throw new Error(`Unknown tool: ${call.name}`);

327 325 

328 const result = await run(JSON.parse(call.arguments));326 const result = await run(JSON.parse(call.arguments));

329 return /** @type {const} */ ({327 return {

330 type: "function_call_output",328 type: "function_call_output",

331 call_id: call.call_id,329 call_id: call.call_id,

332 output: JSON.stringify(result),330 output: JSON.stringify(result),

333 // Preserve caller so the runtime can resume the correct program.331 // Preserve caller so the runtime can resume the correct program.

334 caller: call.caller,332 caller: call.caller,

335 });333 };

336 })334 })

337 );335 );

338 336 

Details

1862import { Agent, run, withTrace, shellTool } from "@openai/agents";1862import { Agent, run, withTrace, shellTool } from "@openai/agents";

1863 1863 

1864class LocalShell {1864class LocalShell {

1865 /** @returns {Promise<import("@openai/agents").ShellResult>} */

1866 async run(action) {1865 async run(action) {

1867 return {1866 return {

1868 output: [1867 output: [

Details

1949#### Example1949#### Example

1950 1950 

1951```python1951```python

1952from typing import Dict, List, Literal1952from typing import Literal

1953 1953 

1954State = Literal["verify", "resolve"]1954State = Literal["verify", "resolve"]

1955 1955 

1956# Allowed transitions1956# Allowed transitions

1957TRANSITIONS: Dict[State, List[State]] = {1957TRANSITIONS: dict[State, list[State]] = {

1958 "verify": ["resolve"],1958 "verify": ["resolve"],

1959 "resolve": [], # terminal1959 "resolve": [], # terminal

1960}1960}


1980 1980 

1981 1981 

1982# Minimal business tools per state1982# Minimal business tools per state

1983TOOLS_BY_STATE: Dict[State, List[dict]] = {1983TOOLS_BY_STATE: dict[State, list[dict]] = {

1984 "verify": [1984 "verify": [

1985 {1985 {

1986 "type": "function",1986 "type": "function",


2011}2011}

2012 2012 

2013# Short, phase-specific instructions2013# Short, phase-specific instructions

2014INSTRUCTIONS_BY_STATE: Dict[State, str] = {2014INSTRUCTIONS_BY_STATE: dict[State, str] = {

2015 "verify": (2015 "verify": (

2016 "# Role & Objective\n"2016 "# Role & Objective\n"

2017 "Verify identity to access the account.\n\n"2017 "Verify identity to access the account.\n\n"

Details

94Use `session.instructions.append` for guardrail steering. It can interrupt speech in progress and apply a new instruction. For example, after your application blocks a request, send:94Use `session.instructions.append` for guardrail steering. It can interrupt speech in progress and apply a new instruction. For example, after your application blocks a request, send:

95 95 

96```javascript96```javascript

97/**

98 * @param {import("openai/resources/live/ws").LiveWS | import("openai/resources/live/sideband/ws").SidebandWS} connection

99 */

100export function sendUpdate(connection) {97export function sendUpdate(connection) {

101 connection.send({98 connection.send({

102 type: "session.instructions.append",99 type: "session.instructions.append",

Details

229audio.controls = true;229audio.controls = true;

230document.body.append(start, stop, status, audio);230document.body.append(start, stop, status, audio);

231 231 

232/** @type {RTCPeerConnection | undefined} */

233let peer;232let peer;

234/** @type {RTCDataChannel | undefined} */233 

235let events;234let events;

236/** @type {MediaStream | undefined} */235 

237let microphone;236let microphone;

238/** @type {ReturnType<typeof setTimeout> | undefined} */237 

239let closeTimeout;238let closeTimeout;

240let ready = false;239let ready = false;

241let finalized = false;240let finalized = false;


273 // Create the event channel before creating the SDP offer.272 // Create the event channel before creating the SDP offer.

274 events = connection.createDataChannel("oai-events");273 events = connection.createDataChannel("oai-events");

275 events.addEventListener("message", ({ data }) => {274 events.addEventListener("message", ({ data }) => {

276 /** @type {import("openai/resources/live/live").ServerEvent} */

277 const event = JSON.parse(data);275 const event = JSON.parse(data);

278 if (event.type === "session.started") {276 if (event.type === "session.started") {

279 ready = true;277 ready = true;


324 body: JSON.stringify({ sdp }),322 body: JSON.stringify({ sdp }),

325 });323 });

326 if (!response.ok) throw new Error(await response.text());324 if (!response.ok) throw new Error(await response.text());

327 /** @type {import("openai/resources/live/live").LiveCreateResponse} */325 

328 const result = await response.json();326 const result = await response.json();

329 console.log("Created session", result.session.id);327 console.log("Created session", result.session.id);

330 await connection.setRemoteDescription({328 await connection.setRemoteDescription({

Details

36let closing = false;36let closing = false;

37let finalized = false;37let finalized = false;

38let pendingByte = Buffer.alloc(0);38let pendingByte = Buffer.alloc(0);

39/** @type {ReturnType<typeof setTimeout> | undefined} */39 

40let closeTimeout;40let closeTimeout;

41 41 

42ws.socket.on("open", () => {42ws.socket.on("open", () => {

Details

93 93 

94```ruby94```ruby

95require "async"95require "async"

96require "async/http/endpoint"96require "openai"

97require "async/websocket/client"

98require "json"

99 97 

100def wait_for_response(connection)98def wait_for_response(connection)

101 while (message = connection.read)99 while (event = connection.receive)

102 event = JSON.parse(message.to_str)100 case event.type.to_s

103 case event.fetch("type")101 when "response.completed" then return event.response

104 when "response.completed" then return event.fetch("response")

105 when "response.failed", "response.incomplete", "error"102 when "response.failed", "response.incomplete", "error"

106 raise "Response failed: #{JSON.generate(event)}"103 raise "Response failed: #{event.to_json}"

107 end104 end

108 end105 end

109 raise "Connection closed before the response finished"106 raise "Connection closed before the response finished"

110end107end

111 108 

112def print_response(response)109client = OpenAI::Client.new

113 response.fetch("output").each do |item|

114 next unless item["type"] == "message"

115 

116 item.fetch("content").each { |part| puts(part.fetch("text")) if part["type"] == "output_text" }

117 end

118end

119 

120endpoint = Async::HTTP::Endpoint.parse("wss://api.openai.com/v1/responses", timeout: 10, alpn_protocols: ["http/1.1"])

121headers = { "Authorization" => "Bearer #{ENV.fetch("OPENAI_API_KEY")}" }

122Sync do |task|110Sync do |task|

123 task.with_timeout(120) do111 task.with_timeout(120) do

124 Async::WebSocket::Client.connect(endpoint, headers: headers) do |connection|112 client.responses.connect(request_options: { timeout: 10 }) do |connection|

125 connection.write(113 connection.response.create(

126 JSON.generate(114 stream_id: "main", model: "gpt-6-astra", store: false,

127 type: "response.create", stream_id: "main", model: "gpt-6-astra", store: false,

128 input: [115 input: [

129 {116 {

130 role: "user",117 role: "user",


132 }119 }

133 ], tools: []120 ], tools: []

134 )121 )

135 )122 puts(wait_for_response(connection).output_text)

136 connection.flush

137 print_response(wait_for_response(connection))

138 end123 end

139 end124 end

140end125end


158 143 

159const client = new OpenAI();144const client = new OpenAI();

160const model = "gpt-6-astra";145const model = "gpt-6-astra";

161/** @type {OpenAI.Responses.FunctionTool[]} */146 

162const tools = [147const tools = [

163 {148 {

164 type: "function",149 type: "function",


169 },154 },

170];155];

171 156 

172/** @param {ResponsesWS} ws */

173async function waitForResponse(ws) {157async function waitForResponse(ws) {

174 for await (const event of ws) {158 for await (const event of ws) {

175 if (event.type === "error") throw event.error;159 if (event.type === "error") throw event.error;


314 298 

315```ruby299```ruby

316require "async"300require "async"

317require "async/http/endpoint"301require "openai"

318require "async/websocket/client"

319require "json"302require "json"

320 303 

321def wait_for_response(connection)304def wait_for_response(connection)

322 while (message = connection.read)305 while (event = connection.receive)

323 event = JSON.parse(message.to_str)306 case event.type.to_s

324 case event.fetch("type")307 when "response.completed" then return event.response

325 when "response.completed" then return event.fetch("response")

326 when "response.failed", "response.incomplete", "error"308 when "response.failed", "response.incomplete", "error"

327 raise "Response failed: #{JSON.generate(event)}"309 raise "Response failed: #{event.to_json}"

328 end310 end

329 end311 end

330 raise "Connection closed before the response finished"312 raise "Connection closed before the response finished"

331end313end

332 314 

333def print_response(response)

334 response.fetch("output").each do |item|

335 next unless item["type"] == "message"

336 

337 item.fetch("content").each { |part| puts(part.fetch("text")) if part["type"] == "output_text" }

338 end

339end

340 

341tools = [315tools = [

342 {316 {

343 type: "function",317 type: "function",


353 }327 }

354]328]

355 329 

356endpoint = Async::HTTP::Endpoint.parse("wss://api.openai.com/v1/responses", timeout: 10, alpn_protocols: ["http/1.1"])330client = OpenAI::Client.new

357headers = { "Authorization" => "Bearer #{ENV.fetch("OPENAI_API_KEY")}" }

358Sync do |task|331Sync do |task|

359 task.with_timeout(120) do332 task.with_timeout(120) do

360 Async::WebSocket::Client.connect(endpoint, headers: headers) do |connection|333 client.responses.connect(request_options: { timeout: 10 }) do |connection|

361 connection.write(334 connection.response.create(

362 JSON.generate(335 stream_id: "main", model: "gpt-6-astra", store: false,

363 type: "response.create", stream_id: "main", model: "gpt-6-astra", store: false,

364 input: "Find the failing test and suggest a fix.", tools: tools,336 input: "Find the failing test and suggest a fix.", tools: tools,

365 tool_choice: {337 tool_choice: {

366 type: "function",338 type: "function",

367 name: "get_test_results"339 name: "get_test_results"

368 }, parallel_tool_calls: false340 }, parallel_tool_calls: false

369 )341 )

370 )

371 connection.flush

372 response = wait_for_response(connection)342 response = wait_for_response(connection)

373 call = response.fetch("output").find { |item| item["type"] == "function_call" }343 call = response.output.grep(OpenAI::Responses::ResponseFunctionToolCall).first

374 unless call && call["name"] == "get_test_results" && JSON.parse(call.fetch("arguments")) == {}344 unless call && call.name == "get_test_results" && JSON.parse(call.arguments) == {}

375 raise "Expected a get_test_results call with no arguments"345 raise "Expected a get_test_results call with no arguments"

376 end346 end

377 347 


380 test: "test_fizz_buzz",350 test: "test_fizz_buzz",

381 failure: 'Expected "FizzBuzz" for 15, got "Fizz".'351 failure: 'Expected "FizzBuzz" for 15, got "Fizz".'

382 }352 }

383 connection.write(353 connection.response.create(

384 JSON.generate(354 stream_id: "main", model: "gpt-6-astra", store: false,

385 type: "response.create", stream_id: "main", model: "gpt-6-astra", store: false,355 previous_response_id: response.id,

386 previous_response_id: response.fetch("id"),

387 input: [356 input: [

388 {357 {

389 type: "function_call_output",358 type: "function_call_output",

390 call_id: call.fetch("call_id"),359 call_id: call.call_id,

391 output: JSON.generate(result)360 output: JSON.generate(result)

392 },361 },

393 {362 {


397 ],366 ],

398 tools: tools, tool_choice: "none"367 tools: tools, tool_choice: "none"

399 )368 )

400 )369 puts(wait_for_response(connection).output_text)

401 connection.flush

402 print_response(wait_for_response(connection))

403 end370 end

404 end371 end

405end372end


525 492 

526```ruby493```ruby

527require "async"494require "async"

528require "async/http/endpoint"

529require "async/websocket/client"

530require "json"

531 

532require "openai"495require "openai"

533 496 

534def wait_for_response(connection)497def wait_for_response(connection)

535 while (message = connection.read)498 while (event = connection.receive)

536 event = JSON.parse(message.to_str)499 case event.type.to_s

537 case event.fetch("type")500 when "response.completed" then return event.response

538 when "response.completed" then return event.fetch("response")

539 when "response.failed", "response.incomplete", "error"501 when "response.failed", "response.incomplete", "error"

540 raise "Response failed: #{JSON.generate(event)}"502 raise "Response failed: #{event.to_json}"

541 end503 end

542 end504 end

543 raise "Connection closed before the response finished"505 raise "Connection closed before the response finished"

544end506end

545 507 

546def print_response(response)

547 response.fetch("output").each do |item|

548 next unless item["type"] == "message"

549 

550 item.fetch("content").each { |part| puts(part.fetch("text")) if part["type"] == "output_text" }

551 end

552end

553 

554client = OpenAI::Client.new508client = OpenAI::Client.new

555compacted = client.responses.compact(509compacted = client.responses.compact(

556 model: "gpt-6-astra",510 model: "gpt-6-astra",


567 content: "Continue from here."521 content: "Continue from here."

568}522}

569 523 

570endpoint = Async::HTTP::Endpoint.parse("wss://api.openai.com/v1/responses", timeout: 10, alpn_protocols: ["http/1.1"])

571headers = { "Authorization" => "Bearer #{ENV.fetch("OPENAI_API_KEY")}" }

572Sync do |task|524Sync do |task|

573 task.with_timeout(120) do525 task.with_timeout(120) do

574 Async::WebSocket::Client.connect(endpoint, headers: headers) do |connection|526 client.responses.connect(request_options: { timeout: 10 }) do |connection|

575 connection.write(527 connection.response.create(

576 JSON.generate(528 stream_id: "main", model: "gpt-6-astra", store: false,

577 type: "response.create", stream_id: "main", model: "gpt-6-astra", store: false,

578 input: next_input, tools: []529 input: next_input, tools: []

579 )530 )

580 )531 puts(wait_for_response(connection).output_text)

581 connection.flush

582 print_response(wait_for_response(connection))

583 end532 end

584 end533 end

585end534end


657 606 

658const client = new OpenAI();607const client = new OpenAI();

659 608 

660/** @type {Map<string, string>} */

661const latestResponseIdByLane = new Map();609const latestResponseIdByLane = new Map();

662 610 

663/**

664 * @param {ResponsesWS} ws

665 * @param {string} streamId

666 * @param {string} text

667 * @param {string} [previousResponseId]

668 */

669function sendCreate(611function sendCreate(

670 ws,612 ws,

671 streamId,613 streamId,


688 });630 });

689}631}

690 632 

691/** @param {ReturnType<ResponsesWS["stream"]>} events */

692async function readMessage(events) {633async function readMessage(events) {

693 while (true) {634 while (true) {

694 const { value: event, done } = await events.next();635 const { value: event, done } = await events.next();


709 }650 }

710}651}

711 652 

712/** @param {ReturnType<ResponsesWS["stream"]>} events @param {Set<string>} expectedStreamIds */

713async function drainUntilComplete(events, expectedStreamIds) {653async function drainUntilComplete(events, expectedStreamIds) {

714 const remaining = new Set(expectedStreamIds);654 const remaining = new Set(expectedStreamIds);

715 while (remaining.size > 0) {655 while (remaining.size > 0) {


723 }663 }

724}664}

725 665 

726/** @param {ReturnType<ResponsesWS["stream"]>} events @param {string} streamId */

727async function waitForInProgress(events, streamId) {666async function waitForInProgress(events, streamId) {

728 while (true) {667 while (true) {

729 const message = await readMessage(events);668 const message = await readMessage(events);


873 812 

874```ruby813```ruby

875require "async"814require "async"

876require "async/http/endpoint"815require "openai"

877require "async/websocket/client"

878require "json"816require "json"

879 817 

880def send_create(connection, stream_id, text, previous_response_id = nil)818def send_create(connection, stream_id, text, previous_response_id = nil)

881 payload = {819 payload = {

882 type: "response.create",

883 stream_id: stream_id,820 stream_id: stream_id,

884 model: "gpt-6-astra",821 model: "gpt-6-astra",

885 store: false,822 store: false,


891 ]828 ]

892 }829 }

893 payload[:previous_response_id] = previous_response_id if previous_response_id830 payload[:previous_response_id] = previous_response_id if previous_response_id

894 connection.write(JSON.generate(payload))831 connection.response.create(**payload)

895 connection.flush

896end832end

897 833 

898def read_event(connection)834def read_event(connection)

899 message = connection.read or raise "Connection closed before all responses finished"835 event = connection.receive or raise "Connection closed before all responses finished"

900 event = JSON.parse(message.to_str)836 if ["response.failed", "response.incomplete", "error"].include?(event.type.to_s)

901 if ["response.failed", "response.incomplete", "error"].include?(event["type"])837 raise "Response failed: #{event.to_json}"

902 raise "Response failed: #{JSON.generate(event)}"

903 end838 end

904 839 

905 event840 event


909 remaining = lanes.dup844 remaining = lanes.dup

910 until remaining.empty?845 until remaining.empty?

911 event = read_event(connection)846 event = read_event(connection)

912 lane = event["stream_id"]847 next unless event.type.to_s == "response.completed"

913 next unless remaining.include?(lane) && event["type"] == "response.completed"

914 848 

915 latest_ids[lane] = event.fetch("response").fetch("id")849 lane = event.stream_id

850 next unless remaining.include?(lane)

851 

852 latest_ids[lane] = event.response.id

916 remaining.delete(lane)853 remaining.delete(lane)

917 end854 end

918end855end

919 856 

920endpoint = Async::HTTP::Endpoint.parse("wss://api.openai.com/v1/responses", timeout: 10, alpn_protocols: ["http/1.1"])857client = OpenAI::Client.new

921headers = { "Authorization" => "Bearer #{ENV.fetch("OPENAI_API_KEY")}" }

922Sync do |task|858Sync do |task|

923 task.with_timeout(120) do859 task.with_timeout(120) do

924 Async::WebSocket::Client.connect(endpoint, headers: headers) do |connection|860 client.responses.connect(request_options: { timeout: 10 }) do |connection|

925 latest_ids = {}861 latest_ids = {}

926 send_create(connection, "planner", "Draft a deployment plan for a stateless API service.")862 send_create(connection, "planner", "Draft a deployment plan for a stateless API service.")

927 send_create(connection, "research", "List common deployment risks for a stateless API service.")863 send_create(connection, "research", "List common deployment risks for a stateless API service.")


931 # Let the fork load its parent before advancing the original lane's cache.867 # Let the fork load its parent before advancing the original lane's cache.

932 loop do868 loop do

933 event = read_event(connection)869 event = read_event(connection)

934 break if event["type"] == "response.in_progress" && event["stream_id"] == "critic"870 break if event.type.to_s == "response.in_progress" && event.stream_id == "critic"

935 end871 end

936 send_create(connection, "planner", "Add rollback and monitoring steps.", parent_id)872 send_create(connection, "planner", "Add rollback and monitoring steps.", parent_id)

937 drain_responses(connection, ["critic", "planner"], latest_ids)873 drain_responses(connection, ["critic", "planner"], latest_ids)

Details

460 460 

461const sts = new STSClient({ region: awsRegion });461const sts = new STSClient({ region: awsRegion });

462 462 

463/** @returns {import("openai/auth/index").SubjectTokenProvider} */

464function awsOutboundWebIdentityTokenProvider() {463function awsOutboundWebIdentityTokenProvider() {

465 return {464 return {

466 tokenType: "jwt",465 tokenType: "jwt",


1230 );1229 );

1231}1230}

1232 1231 

1233/** @returns {import("openai/auth/index").SubjectTokenProvider} */

1234function mountedEksServiceAccountTokenProvider(path) {1232function mountedEksServiceAccountTokenProvider(path) {

1235 return {1233 return {

1236 tokenType: "jwt",1234 tokenType: "jwt",

Details

469 );469 );

470}470}

471 471 

472/** @returns {import("openai/auth/index").SubjectTokenProvider} */

473function githubActionsOIDCTokenProvider(requestURL, requestToken, audience) {472function githubActionsOIDCTokenProvider(requestURL, requestToken, audience) {

474 return {473 return {

475 tokenType: "jwt",474 tokenType: "jwt",

Details

413 );413 );

414}414}

415 415 

416/** @returns {import("openai/auth/index").SubjectTokenProvider} */

417function googleMetadataIdentityTokenProvider(audience) {416function googleMetadataIdentityTokenProvider(audience) {

418 return {417 return {

419 tokenType: "jwt",418 tokenType: "jwt",


1255 );1254 );

1256}1255}

1257 1256 

1258/** @returns {import("openai/auth/index").SubjectTokenProvider} */

1259function mountedGkeServiceAccountTokenProvider(path) {1257function mountedGkeServiceAccountTokenProvider(path) {

1260 return {1258 return {

1261 tokenType: "jwt",1259 tokenType: "jwt",

Details

439 );439 );

440}440}

441 441 

442/** @returns {import("openai/auth/index").SubjectTokenProvider} */

443function mountedServiceAccountTokenProvider(path) {442function mountedServiceAccountTokenProvider(path) {

444 return {443 return {

445 tokenType: "jwt",444 tokenType: "jwt",

Details

429 );429 );

430}430}

431 431 

432/** @returns {import("openai/auth/index").SubjectTokenProvider} */

433function azureManagedIdentityTokenProvider(resource) {432function azureManagedIdentityTokenProvider(resource) {

434 return {433 return {

435 tokenType: "jwt",434 tokenType: "jwt",


1300 );1299 );

1301}1300}

1302 1301 

1303/** @returns {import("openai/auth/index").SubjectTokenProvider} */

1304function mountedAksServiceAccountTokenProvider(path) {1302function mountedAksServiceAccountTokenProvider(path) {

1305 return {1303 return {

1306 tokenType: "jwt",1304 tokenType: "jwt",

Details

499 );499 );

500}500}

501 501 

502/** @returns {import("openai/auth/index").SubjectTokenProvider} */

503function spiffeJwtSvidProvider(path) {502function spiffeJwtSvidProvider(path) {

504 return {503 return {

505 tokenType: "jwt",504 tokenType: "jwt",

quickstart.md +0 −1

Details

1542import OpenAI from "openai";1542import OpenAI from "openai";

1543const client = new OpenAI();1543const client = new OpenAI();

1544 1544 

1545/** @type {OpenAI.Responses.Tool[]} */

1546const tools = [1545const tools = [

1547 {1546 {

1548 type: "function",1547 type: "function",