SpyBara
Go Premium

Documentation 2026-08-20 23:59 UTC to 2026-08-21 19:59 UTC

6 files changed +71 −21. View all changes and history on the product overview
2026
Mon 31 23:00 Fri 28 18:00 Thu 27 22:01 Wed 26 22:57 Tue 25 18:59 Mon 24 22:00 Fri 21 19:59 Thu 20 23:59 Wed 19 18:02 Tue 18 04:58 Mon 17 22:57 Sat 15 01:01 Fri 14 20:01 Thu 13 22:00 Wed 12 02:57 Tue 11 19:59 Mon 10 19:00 Fri 7 00:58 Thu 6 21:58 Wed 5 18:01 Tue 4 22:59 Mon 3 18:01
Details

10 `store=false`. Response data is temporarily stored to disk for roughly 1010 `store=false`. Response data is temporarily stored to disk for roughly 10

11 minutes to enable asynchronous execution and polling.11 minutes to enable asynchronous execution and polling.

12 12 

13For projects using [Modified Abuse

14Monitoring](https://developers.openai.com/api/docs/guides/your-data#modified-abuse-monitoring), including

15enhanced Modified Abuse Monitoring, foreground requests follow standard

16retention when `store` is omitted or set to `true`. Background responses are

17retained after the polling period only when `store=true` is explicitly provided.

18If `store` is omitted or set to `false` for a background request, the response

19is deleted after roughly 10 minutes.

20 

13Generate a response in the background21Generate a response in the background

14 22 

15```bash23```bash

Details

158 158 

159Fast mode charges a per-token premium compared with Standard processing. All processing modes count toward your annual Enterprise spend commitment, and eligible cached input tokens receive the same discounts available for Standard processing.159Fast mode charges a per-token premium compared with Standard processing. All processing modes count toward your annual Enterprise spend commitment, and eligible cached input tokens receive the same discounts available for Standard processing.

160 160 

161For GPT-5.6 Sol, Fast mode costs twice the corresponding Standard rate. Short-context requests cost $8 per 1 million input tokens and $40 per 1 million output tokens; long-context requests cost $16 per 1 million input tokens and $60 per 1 million output tokens. GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026. See [pricing details](https://developers.openai.com/api/docs/pricing?latest-pricing=fast).

162 

161To review usage, open the usage dashboard, select Responses or Chat Completions, and group by service tier. To review costs, group by line item.163To review usage, open the usage dashboard, select Responses or Chat Completions, and group by service tier. To review costs, group by line item.

162 164 

163### Which models and modalities support Fast mode?165### Which models and modalities support Fast mode?

Details

133### Additional differences133### Additional differences

134 134 

135- Responses are stored by default. Chat completions are stored by default for new accounts. To disable storage in either API, set `store: false`.135- Responses are stored by default. Chat completions are stored by default for new accounts. To disable storage in either API, set `store: false`.

136- [Reasoning](https://developers.openai.com/api/docs/guides/reasoning) models have a richer experience in the Responses API with [improved tool usage](https://developers.openai.com/api/docs/guides/reasoning#keeping-reasoning-items-in-context). Starting with GPT-5.4, tool calling is not supported in Chat Completions with `reasoning: none`.136- [Reasoning](https://developers.openai.com/api/docs/guides/reasoning) models have a richer experience in the Responses API with [improved tool usage](https://developers.openai.com/api/docs/guides/reasoning#keeping-reasoning-items-in-context). Starting with GPT-5.4, Chat Completions does not support tool calling with `reasoning_effort` values other than `none`.

137- Structured Outputs API shape is different. Instead of `response_format`, use `text.format` in Responses. Learn more in the [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) guide.137- Structured Outputs API shape is different. Instead of `response_format`, use `text.format` in Responses. Learn more in the [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) guide.

138- The function-calling API shape is different, both for the function config on the request, and function calls sent back in the response. See the full difference in the [function calling guide](https://developers.openai.com/api/docs/guides/function-calling).138- The function-calling API shape is different, both for the function config on the request, and function calls sent back in the response. See the full difference in the [function calling guide](https://developers.openai.com/api/docs/guides/function-calling).

139- The Responses SDK has an `output_text` helper, which the Chat Completions SDK does not have.139- The Responses SDK has an `output_text` helper, which the Chat Completions SDK does not have.

Details

91 91 

92#### `/v1/responses`92#### `/v1/responses`

93 93 

94- The Responses API has a 30 day Application State retention period by default, or when the `store` parameter is set to `true`. Response data will be stored for at least 30 days.94- Except as noted below, the Responses API has a 30 day Application State retention period by default, or when the `store` parameter is set to `true`. In those cases, response data will be stored for at least 30 days.

95- When Zero Data Retention is enabled for an organization, the `store` parameter will always be treated as `false`, even if the request attempts to set the value to `true`.95- When Zero Data Retention is enabled for an organization, the `store` parameter will always be treated as `false`, even if the request attempts to set the value to `true`.

96- Background mode stores response data to disk for roughly 10 minutes to enable polling.96- Background mode stores response data to disk for roughly 10 minutes to enable polling. For projects using [Modified Abuse Monitoring](#modified-abuse-monitoring), including enhanced Modified Abuse Monitoring, foreground requests follow standard retention when `store` is omitted or set to `true`. Background responses follow the standard retention period only when the request explicitly sets `store=true`. If `store` is omitted or set to `false` for a background request, the response is deleted after the temporary polling period.

97- Audio outputs application state is stored for 1 hour to enable [multi-turn conversations](https://developers.openai.com/api/docs/guides/audio).97- Audio outputs application state is stored for 1 hour to enable [multi-turn conversations](https://developers.openai.com/api/docs/guides/audio).

98- See [image and file inputs](#image-and-file-inputs).98- See [image and file inputs](#image-and-file-inputs).

99- MCP servers (used with the [remote MCP server tool](https://developers.openai.com/api/docs/guides/tools-connectors-mcp)) are third-party services, and data sent to an MCP server is subject to their data retention policies.99- MCP servers (used with the [remote MCP server tool](https://developers.openai.com/api/docs/guides/tools-connectors-mcp)) are third-party services, and data sent to an MCP server is subject to their data retention policies.


166 166 

167For requests to projects with data residency configured, add the domain prefix as defined in the table below to each request.167For requests to projects with data residency configured, add the domain prefix as defined in the table below to each request.

168 168 

169#### Select a processing region per request

170 

171As an alternative to creating a region-specific project, you can select regional processing for an individual request by using the prefixed domain with an API key from a project having Global geography.

172 

173Existing eligibility and data retention control requirements still apply. The selected endpoint and model must also support regional processing, as shown in the table below.

174 

175The following example reuses one client and an API key from a Global project for global, US, and EU requests:

176 

177```python

178from openai import OpenAI

179 

180client = OpenAI()

181 

182# No processing constraint.

183response = client.responses.create(

184 model="gpt-5.6-terra",

185 input="Reply with OK.",

186)

187print(response.output_text)

188 

189# US processing and storage.

190response = client.with_options(

191 base_url="https://us.api.openai.com/v1",

192).responses.create(

193 model="gpt-5.6-terra",

194 input="Reply with OK.",

195)

196print(response.output_text)

197 

198# EU processing and storage.

199response = client.with_options(

200 base_url="https://eu.api.openai.com/v1",

201).responses.create(

202 model="gpt-5.6-terra",

203 input="Reply with OK.",

204)

205print(response.output_text)

206```

207 

208 

169### Which models and features are eligible for data residency?209### Which models and features are eligible for data residency?

170 210 

171The following models and API services are eligible for data residency today for the regions specified below.211The following models and API services are eligible for data residency today for the regions specified below.

models.md +9 −9

Details

29- [davinci-002](/api/docs/models/davinci-002.md): Replacement for the GPT-3 curie and davinci base models29- [davinci-002](/api/docs/models/davinci-002.md): Replacement for the GPT-3 curie and davinci base models

30- [Daybreak Blue](/api/docs/models/daybreak-blue-latest.md): An alias for frontier general-purpose models with safeguards for defensive cybersecurity work.30- [Daybreak Blue](/api/docs/models/daybreak-blue-latest.md): An alias for frontier general-purpose models with safeguards for defensive cybersecurity work.

31- [Daybreak Red](/api/docs/models/daybreak-red-latest.md): An alias for advanced cybersecurity models for authorized vulnerability research and security testing.31- [Daybreak Red](/api/docs/models/daybreak-red-latest.md): An alias for advanced cybersecurity models for authorized vulnerability research and security testing.

32- [GPT Image 1](/api/docs/models/gpt-image-1.md): Our previous image generation model

33- [GPT Image 1.5](/api/docs/models/gpt-image-1.5.md): Our previous image generation model

34- [GPT Image 2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model

35- [GPT Live Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription

36- [GPT Transcribe](/api/docs/models/gpt-transcribe.md): High-accuracy speech-to-text model for file and Realtime input transcription

37- [GPT-3.5 Turbo](/api/docs/models/gpt-3.5-turbo.md): Legacy GPT model for cheaper chat and non-chat tasks32- [GPT-3.5 Turbo](/api/docs/models/gpt-3.5-turbo.md): Legacy GPT model for cheaper chat and non-chat tasks

38- [GPT-4](/api/docs/models/gpt-4.md): An older high-intelligence GPT model33- [GPT-4](/api/docs/models/gpt-4.md): An older high-intelligence GPT model

39- [GPT-4 Turbo](/api/docs/models/gpt-4-turbo.md): An older high-intelligence GPT model34- [GPT-4 Turbo](/api/docs/models/gpt-4-turbo.md): An older high-intelligence GPT model


81- [GPT-5.6 Luna](/api/docs/models/gpt-5.6-luna.md): GPT-5.6 model optimized for cost-sensitive workloads76- [GPT-5.6 Luna](/api/docs/models/gpt-5.6-luna.md): GPT-5.6 model optimized for cost-sensitive workloads

82- [GPT-5.6 Sol](/api/docs/models/gpt-5.6-sol.md): Frontier model for complex professional work77- [GPT-5.6 Sol](/api/docs/models/gpt-5.6-sol.md): Frontier model for complex professional work

83- [GPT-5.6 Terra](/api/docs/models/gpt-5.6-terra.md): GPT-5.6 model that balances intelligence and cost78- [GPT-5.6 Terra](/api/docs/models/gpt-5.6-terra.md): GPT-5.6 model that balances intelligence and cost

84- [gpt-audio](/api/docs/models/gpt-audio.md): For audio inputs and outputs with Chat Completions API79- [GPT-Audio](/api/docs/models/gpt-audio.md): For audio inputs and outputs with Chat Completions API

85- [gpt-audio-1.5](/api/docs/models/gpt-audio-1.5.md): The best voice model for audio in, audio out with Chat Completions.80- [GPT-Audio mini](/api/docs/models/gpt-audio-mini.md): A cost-efficient version of GPT Audio

86- [gpt-audio-mini](/api/docs/models/gpt-audio-mini.md): A cost-efficient version of GPT Audio81- [GPT-Audio-1.5](/api/docs/models/gpt-audio-1.5.md): The best voice model for audio in, audio out with Chat Completions.

87- [gpt-image-1-mini](/api/docs/models/gpt-image-1-mini.md): A cost-efficient version of GPT Image 182- [GPT-Image-1](/api/docs/models/gpt-image-1.md): Our previous image generation model

83- [GPT-Image-1 mini](/api/docs/models/gpt-image-1-mini.md): A cost-efficient version of GPT Image 1

84- [GPT-Image-1.5](/api/docs/models/gpt-image-1.5.md): Our previous image generation model

85- [GPT-Image-2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model

86- [GPT-Live-Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription

88- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU87- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU

89- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency88- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency

90- [GPT-Realtime](/api/docs/models/gpt-realtime.md): Model capable of realtime text and audio inputs and outputs89- [GPT-Realtime](/api/docs/models/gpt-realtime.md): Model capable of realtime text and audio inputs and outputs


95- [GPT-Realtime-2.1 mini](/api/docs/models/gpt-realtime-2.1-mini.md): Reasoning model with tool use94- [GPT-Realtime-2.1 mini](/api/docs/models/gpt-realtime-2.1-mini.md): Reasoning model with tool use

96- [GPT-Realtime-Translate](/api/docs/models/gpt-realtime-translate.md): Streaming speech-to-speech translation model95- [GPT-Realtime-Translate](/api/docs/models/gpt-realtime-translate.md): Streaming speech-to-speech translation model

97- [GPT-Realtime-Whisper](/api/docs/models/gpt-realtime-whisper.md): Streaming speech-to-text model for realtime transcription96- [GPT-Realtime-Whisper](/api/docs/models/gpt-realtime-whisper.md): Streaming speech-to-text model for realtime transcription

97- [GPT-Transcribe](/api/docs/models/gpt-transcribe.md): High-accuracy speech-to-text model for file and Realtime input transcription

98- [o1](/api/docs/models/o1.md): Previous full o-series reasoning model98- [o1](/api/docs/models/o1.md): Previous full o-series reasoning model

99- [o1 Preview](/api/docs/models/o1-preview.md): Preview of our first o-series reasoning model99- [o1 Preview](/api/docs/models/o1-preview.md): Preview of our first o-series reasoning model

100- [o1-mini](/api/docs/models/o1-mini.md): A small model alternative to o1100- [o1-mini](/api/docs/models/o1-mini.md): A small model alternative to o1

models/all.md +9 −9

Details

29- [davinci-002](/api/docs/models/davinci-002.md): Replacement for the GPT-3 curie and davinci base models29- [davinci-002](/api/docs/models/davinci-002.md): Replacement for the GPT-3 curie and davinci base models

30- [Daybreak Blue](/api/docs/models/daybreak-blue-latest.md): An alias for frontier general-purpose models with safeguards for defensive cybersecurity work.30- [Daybreak Blue](/api/docs/models/daybreak-blue-latest.md): An alias for frontier general-purpose models with safeguards for defensive cybersecurity work.

31- [Daybreak Red](/api/docs/models/daybreak-red-latest.md): An alias for advanced cybersecurity models for authorized vulnerability research and security testing.31- [Daybreak Red](/api/docs/models/daybreak-red-latest.md): An alias for advanced cybersecurity models for authorized vulnerability research and security testing.

32- [GPT Image 1](/api/docs/models/gpt-image-1.md): Our previous image generation model

33- [GPT Image 1.5](/api/docs/models/gpt-image-1.5.md): Our previous image generation model

34- [GPT Image 2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model

35- [GPT Live Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription

36- [GPT Transcribe](/api/docs/models/gpt-transcribe.md): High-accuracy speech-to-text model for file and Realtime input transcription

37- [GPT-3.5 Turbo](/api/docs/models/gpt-3.5-turbo.md): Legacy GPT model for cheaper chat and non-chat tasks32- [GPT-3.5 Turbo](/api/docs/models/gpt-3.5-turbo.md): Legacy GPT model for cheaper chat and non-chat tasks

38- [GPT-4](/api/docs/models/gpt-4.md): An older high-intelligence GPT model33- [GPT-4](/api/docs/models/gpt-4.md): An older high-intelligence GPT model

39- [GPT-4 Turbo](/api/docs/models/gpt-4-turbo.md): An older high-intelligence GPT model34- [GPT-4 Turbo](/api/docs/models/gpt-4-turbo.md): An older high-intelligence GPT model


81- [GPT-5.6 Luna](/api/docs/models/gpt-5.6-luna.md): GPT-5.6 model optimized for cost-sensitive workloads76- [GPT-5.6 Luna](/api/docs/models/gpt-5.6-luna.md): GPT-5.6 model optimized for cost-sensitive workloads

82- [GPT-5.6 Sol](/api/docs/models/gpt-5.6-sol.md): Frontier model for complex professional work77- [GPT-5.6 Sol](/api/docs/models/gpt-5.6-sol.md): Frontier model for complex professional work

83- [GPT-5.6 Terra](/api/docs/models/gpt-5.6-terra.md): GPT-5.6 model that balances intelligence and cost78- [GPT-5.6 Terra](/api/docs/models/gpt-5.6-terra.md): GPT-5.6 model that balances intelligence and cost

84- [gpt-audio](/api/docs/models/gpt-audio.md): For audio inputs and outputs with Chat Completions API79- [GPT-Audio](/api/docs/models/gpt-audio.md): For audio inputs and outputs with Chat Completions API

85- [gpt-audio-1.5](/api/docs/models/gpt-audio-1.5.md): The best voice model for audio in, audio out with Chat Completions.80- [GPT-Audio mini](/api/docs/models/gpt-audio-mini.md): A cost-efficient version of GPT Audio

86- [gpt-audio-mini](/api/docs/models/gpt-audio-mini.md): A cost-efficient version of GPT Audio81- [GPT-Audio-1.5](/api/docs/models/gpt-audio-1.5.md): The best voice model for audio in, audio out with Chat Completions.

87- [gpt-image-1-mini](/api/docs/models/gpt-image-1-mini.md): A cost-efficient version of GPT Image 182- [GPT-Image-1](/api/docs/models/gpt-image-1.md): Our previous image generation model

83- [GPT-Image-1 mini](/api/docs/models/gpt-image-1-mini.md): A cost-efficient version of GPT Image 1

84- [GPT-Image-1.5](/api/docs/models/gpt-image-1.5.md): Our previous image generation model

85- [GPT-Image-2](/api/docs/models/gpt-image-2.md): State-of-the-art image generation model

86- [GPT-Live-Transcribe](/api/docs/models/gpt-live-transcribe.md): Low-latency speech-to-text model for realtime transcription

88- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU87- [gpt-oss-120b](/api/docs/models/gpt-oss-120b.md): Most powerful open-weight model, fits into an H100 GPU

89- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency88- [gpt-oss-20b](/api/docs/models/gpt-oss-20b.md): Medium-sized open-weight model for low latency

90- [GPT-Realtime](/api/docs/models/gpt-realtime.md): Model capable of realtime text and audio inputs and outputs89- [GPT-Realtime](/api/docs/models/gpt-realtime.md): Model capable of realtime text and audio inputs and outputs


95- [GPT-Realtime-2.1 mini](/api/docs/models/gpt-realtime-2.1-mini.md): Reasoning model with tool use94- [GPT-Realtime-2.1 mini](/api/docs/models/gpt-realtime-2.1-mini.md): Reasoning model with tool use

96- [GPT-Realtime-Translate](/api/docs/models/gpt-realtime-translate.md): Streaming speech-to-speech translation model95- [GPT-Realtime-Translate](/api/docs/models/gpt-realtime-translate.md): Streaming speech-to-speech translation model

97- [GPT-Realtime-Whisper](/api/docs/models/gpt-realtime-whisper.md): Streaming speech-to-text model for realtime transcription96- [GPT-Realtime-Whisper](/api/docs/models/gpt-realtime-whisper.md): Streaming speech-to-text model for realtime transcription

97- [GPT-Transcribe](/api/docs/models/gpt-transcribe.md): High-accuracy speech-to-text model for file and Realtime input transcription

98- [o1](/api/docs/models/o1.md): Previous full o-series reasoning model98- [o1](/api/docs/models/o1.md): Previous full o-series reasoning model

99- [o1 Preview](/api/docs/models/o1-preview.md): Preview of our first o-series reasoning model99- [o1 Preview](/api/docs/models/o1-preview.md): Preview of our first o-series reasoning model

100- [o1-mini](/api/docs/models/o1-mini.md): A small model alternative to o1100- [o1-mini](/api/docs/models/o1-mini.md): A small model alternative to o1