1---1---
2latestModelInfo:2latestModelInfo:
3 model: gpt-5.43 model: gpt-5.5
4 migrationGuide: /api/docs/guides/upgrading-to-gpt-5p4.md4 migrationGuide: /api/docs/guides/upgrading-to-gpt-5p5.md
5 promptingGuide: /api/docs/guides/prompt-guidance.md5 promptingGuide: /api/docs/guides/prompt-guidance.md
6---6---
7 7
8# Using GPT-5.48# Using GPT-5.5
9 9
10import {10## Introduction
11 Bolt,
12 Code,
13 Sparkles,
14} from "@components/react/oai/platform/ui/Icon.react";
15 11
12GPT-5.5 raises the baseline for complex production workflows. It’s a strong fit for coding use cases, tool-heavy agents, grounded assistants, long-context retrieval, product-spec-to-plan workflows, and customer-facing workflows where execution quality and response polish are critical.
16 13
14To get the most out of GPT-5.5, treat it as a new model family to tune for, not a drop-in replacement for `gpt-5.2` or `gpt-5.4`. Begin migration with a fresh baseline instead of carrying over every instruction from an older prompt stack. Start with the smallest prompt that preserves the product contract, then tune reasoning effort, verbosity, tool descriptions, and output format against representative examples.
17 15
16GPT-5.5 supports all API features that were already available with GPT-5.4, including [prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching), [hosted tools](https://developers.openai.com/api/docs/guides/tools#available-tools), [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search), [compaction](https://developers.openai.com/api/docs/guides/compaction), and `phase` handling for manually replayed assistant items.
18 17
19export const fastResponses = <>18See the [GPT-5.5 Prompting Guide](https://developers.openai.com/api/docs/guides/prompt-guidance?model=gpt-5.5) for examples of successful prompting patterns.
20 19
21GPT-5.4 has a new reasoning mode: `none` for low-latency interactions. By default, GPT-5.4 reasoning is set to `none`.20## What's new
22 21
23<br />22- **More efficient reasoning:** GPT-5.5 reaches strong results with fewer reasoning tokens than prior models, even at the same reasoning effort. This is especially useful in complex, tool-heavy, or multi-step workflows where token savings compound.
24<br />23- **Stronger task execution with outcome-first prompts:** GPT-5.5 is better at working from a clear goal, preserving constraints, and turning product intent into concrete next steps. Describe the expected outcome, success criteria, allowed side effects, evidence rules, and output shape. Avoid step-by-step process guidance unless the exact path matters.
24- **Stronger and more precise tool use:** GPT-5.5 is especially useful on large tool surfaces, multi-step service workflows, and long-running agent tasks. It tends to be more precise in tool selection and argument use.
25- **Tone is often more polished, but can be more direct:** GPT-5.5 often produces warmer, more readable answers with less prompt scaffolding.
25 26
26This behavior will more closely (but not exactly!) match non-reasoning models like 27## Behavioral changes
27 28
28[GPT-4.1](https://developers.openai.com/api/docs/models/gpt-4.1). We expect GPT-5.4 to produce291. **Reasoning effort now defaults to `medium`:** GPT-5.5 defaults to `medium` reasoning effort. Treat `medium` as the recommended balanced starting point for quality, reliability, latency, and cost. For latency-sensitive workflows, evaluate `low` before `none` when tool use, planning, search, or multi-step decision making still matters. Reserve `none` for latency-critical tasks that don't need reasoning or multi-chained tool calls, such as lightweight voice turns, fast information retrieval, and classification. Increase to `high` or `xhigh` only when evals show a measurable quality gain that justifies the extra latency and cost. See the [Reasoning models documentation](https://developers.openai.com/api/docs/guides/reasoning) for more details on recommended settings.
29more intelligent responses than GPT-4.1, but when speed and maximum context
30length are paramount, you might consider using GPT-4.1 instead.
31 30
32Fast, low latency response options31 Higher reasoning effort isn't automatically better. If the task has conflicting instructions, weak stopping criteria, or open-ended tool access, higher effort can lead to overthinking, unnecessary searching, or output quality regressions. Increase effort only when evals show a measurable quality gain.
33 32
34```javascript332. **Image inputs preserve more visual detail by default:** GPT-5.5 updates the default handling for image inputs to preserve more visual detail and improve computer use performance. When `image_detail` is unset or set to `auto`, the model now uses `original` behavior, preserving images without resizing up to 10,240,000 pixels or a 6,000-pixel dimension limit. For `high`, specify the value directly; it preserves images without resizing up to 2,500,000 pixels or a 2,048-pixel dimension limit. `low` now focuses on context efficiency and resizes images above a 512-pixel dimension limit more aggressively than previous models. See the [Images and vision documentation](https://developers.openai.com/api/docs/guides/images-vision).
35import OpenAI from "openai";
36const openai = new OpenAI();
37 34
38const result = await openai.responses.create({353. **Improved instruction following:** GPT-5.5 interprets prompts in a literal and thorough manner, enabling specific, descriptive instructions when the product requires them. Define success criteria and stopping rules, especially for long-running, tool-heavy, or evidence-gathering workflows. See [Write outcome-first prompts](https://developers.openai.com/api/docs/guides/prompt-guidance?model=gpt-5.5#outcome-first-prompts-and-stopping-conditions) and [Keep the right specificity](https://developers.openai.com/api/docs/guides/prompt-guidance?model=gpt-5.5#formatting).
39 model: "gpt-5.4",
40 input: "Write a haiku about code.",
41 reasoning: { effort: "low" },
42 text: { verbosity: "low" },
43});
44 36
45console.log(result.output_text);374. **Default style is more concise and direct:** GPT-5.5 tends to be efficient, direct, and task-oriented by default. This is useful for many production workflows, but customer-facing or conversational experiences may need explicit personality, warmth, rationale, and formatting guidance. Use `text.verbosity` intentionally: `medium` is the default, and `low` is often a better starting point for concise responses. See the [GPT-5.5 prompting guide](https://developers.openai.com/api/docs/guides/prompt-guidance?model=gpt-5.5).
46```
47
48```python
49from openai import OpenAI
50client = OpenAI()
51
52result = client.responses.create(
53 model="gpt-5.4",
54 input="Write a haiku about code.",
55 reasoning={ "effort": "low" },
56 text={ "verbosity": "low" },
57)
58
59print(result.output_text)
60```
61
62```bash
63curl https://api.openai.com/v1/responses \\
64 -H "Content-Type: application/json" \\
65 -H "Authorization: Bearer $OPENAI_API_KEY" \\
66 -d '{
67 "model": "gpt-5.4",
68 "input": "Write a haiku about code.",
69 "reasoning": { "effort": "low" }
70 }'
71```
72
73
74</>;
75
76export const goodResponses = <>
77
78GPT-5.4 is great at reasoning through complex tasks. <strong>For complex tasks like coding and multi-step planning,
79use high reasoning effort.</strong>
80
81<br />
82<br />
83
84Use these configurations when replacing tasks you might have used o3 to tackle.
85We expect GPT-5.4 to produce better results than o3 and o4-mini under most circumstances.
86
87Slower, high reasoning tasks
88
89```javascript
90import OpenAI from "openai";
91const openai = new OpenAI();
92
93const result = await openai.responses.create({
94 model: "gpt-5.4",
95 input: "Find the null pointer exception: ...your code here...",
96 reasoning: { effort: "high" },
97});
98
99console.log(result.output_text);
100```
101
102```python
103from openai import OpenAI
104client = OpenAI()
105
106result = client.responses.create(
107 model="gpt-5.4",
108 input="Find the null pointer exception: ...your code here...",
109 reasoning={ "effort": "high" },
110)
111
112print(result.output_text)
113```
114
115```bash
116curl https://api.openai.com/v1/responses \\
117 -H "Content-Type: application/json" \\
118 -H "Authorization: Bearer $OPENAI_API_KEY" \\
119 -d '{
120 "model": "gpt-5.4",
121 "input": "Find the null pointer exception: ...your code here...",
122 "reasoning": { "effort": "high" }
123 }'
124```
125
126
127</>;
128
129[GPT-5.4](https://developers.openai.com/api/docs/models/gpt-5.4) is our most capable frontier model yet, delivering higher-quality outputs with fewer iterations across ChatGPT, the API, and Codex. It helps people and teams analyze complex information, build production software, and automate multi-step workflows.
130
131GPT-5.5 is currently available in ChatGPT and Codex, with API availability
132 coming soon.
133
134In practice, `gpt-5.4` is the default model for both broad general-purpose work and most coding tasks. Start there when you want one model that can move between software engineering, reasoning, writing, and tool use in the same workflow.
135
136This guide covers key features of the GPT-5 model family and how to get the most out of GPT-5.4.
137
138## Key improvements
139
140Compared with the previous GPT-5.2 model, GPT-5.4 shows improvements in:
141
142- Coding, document understanding, tool use, and instruction following
143- Image perception and multimodal tasks
144- Long-running task execution and multi-step agent workflows
145- Token efficiency and end-to-end performance on tool-heavy workloads
146- Agentic web search and multi-source synthesis, especially for hard-to-locate information
147- Document-heavy and spreadsheet-heavy business workflows in customer service, analytics, and finance
148
149GPT-5.4 brings the coding capabilities of GPT-5.3-Codex to our flagship frontier model. Developers can generate production-quality code, build polished front-end UI, follow repo-specific patterns, and handle multi-file changes with fewer retries. It also has a strong out-of-the-box coding personality, so teams spend less time on prompt tuning.
150
151For agentic workloads, GPT-5.4 reduces end-to-end time across multi-step trajectories and often completes tasks with fewer tokens and tool calls. This makes agents more responsive and lowers the cost of operating complex workflows at scale in the API and Codex.
152
153### New features in GPT-5.4
154
155Like earlier GPT-5 models, GPT-5.4 supports custom tools, parameters to control verbosity and reasoning, and an allowed tools list. GPT-5.4 also introduces several capabilities that make it easier to build powerful agent systems, operate over larger bodies of information, and run more reliable automated workflows:
156
157- **`tool_search` in the API:** GPT-5.4 improves tool search for larger tool ecosystems by using deferred tool loading. This makes tools searchable, loads only the relevant definitions, reduces token usage, and improves tool selection accuracy in real deployments. Learn more in the [tool search guide](https://developers.openai.com/api/docs/guides/tools-tool-search).
158- **1M token context window:** GPT-5.4 supports up to a 1M token context window, making it easier to analyze entire codebases, long document collections, or extended agent trajectories in a single request. Read more in the [1M context window](#1m-context-window) section.
159- **Built-in computer use:** GPT-5.4 is the first mainline model with built-in computer-use capabilities, enabling agents to interact directly with software to complete, verify, and fix tasks in a build-run-verify-fix loop. Learn more in the [computer use guide](https://developers.openai.com/api/docs/guides/tools-computer-use).
160- **Native compaction support:** GPT-5.4 is the first mainline model trained to support compaction, enabling longer agent trajectories while preserving key context.
161
162## Meet the models
163
164In general, `gpt-5.4` is the default model for your most important work across both general-purpose tasks and coding. It replaces the previous `gpt-5.2` model in the API, and `gpt-5.3-codex` in Codex. The model powering ChatGPT is `gpt-5-chat-latest`. For more difficult problems, `gpt-5.4-pro` uses more compute to think longer and provide consistently better answers.
165
166For smaller, faster variants, start with `gpt-5.4-mini` or `gpt-5.4-nano`.
167
168To help you pick the model that best fits your use case, consider these tradeoffs:
169
170| Variant | Best for |
171| ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
172| [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) | General-purpose work, including complex reasoning, broad world knowledge, and code-heavy or multi-step agentic tasks |
173| [`gpt-5.4-pro`](https://developers.openai.com/api/docs/models/gpt-5.4-pro) | Tough problems that may take longer to solve and need deeper reasoning |
174| [`gpt-5.4-mini`](https://developers.openai.com/api/docs/models/gpt-5.4-mini) | High-volume coding, computer use, and agent workflows that still need strong reasoning |
175| [`gpt-5.4-nano`](https://developers.openai.com/api/docs/models/gpt-5.4-nano) | Simple high-throughput tasks where speed and cost matter most |
176
177### Lower reasoning effort
178
179The `reasoning.effort` parameter controls how many reasoning tokens the model generates before producing a response. Earlier reasoning models like o3 supported only `low`, `medium`, and `high`: `low` favored speed and fewer tokens, while `high` favored more thorough reasoning.
180
181Starting with GPT-5.2, the lowest setting is `none` to provide lower-latency interactions. This is the default setting in GPT-5.2 and newer models. If you need more thinking, slowly increase to `medium` and experiment with results.
182
183With reasoning effort set to `none`, prompting is important. To improve the model's reasoning quality, even with the default settings, encourage it to “think” or outline its steps before answering.
184
185Reasoning effort set to none
186
187```bash
188curl --request POST \
189 --url https://api.openai.com/v1/responses \
190 --header "Authorization: Bearer $OPENAI_API_KEY" \
191 --header 'Content-type: application/json' \
192 --data '{
193 "model": "gpt-5.4",
194 "input": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
195 "reasoning": {
196 "effort": "none"
197 }
198}'
199```
200
201```javascript
202import OpenAI from "openai";
203const openai = new OpenAI();
204
205const response = await openai.responses.create({
206 model: "gpt-5.4",
207 input: "How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
208 reasoning: {
209 effort: "none"
210 }
211});
212
213console.log(response);
214```
215
216```python
217from openai import OpenAI
218client = OpenAI()
219
220response = client.responses.create(
221 model="gpt-5.4",
222 input="How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
223 reasoning={
224 "effort": "none"
225 }
226)
227
228print(response)
229```
230
231
232### Verbosity
233
234Verbosity determines how many output tokens are generated. Lowering the number of tokens reduces overall latency. While the model's reasoning approach stays mostly the same, the model finds ways to answer more concisely—which can either improve or diminish answer quality, depending on your use case. Here are some scenarios for both ends of the verbosity spectrum:
235
236- **High verbosity:** Use when you need the model to provide thorough explanations of documents or perform extensive code refactoring.
237- **Low verbosity:** Best for situations where you want concise answers or simple code generation, such as SQL queries.
238
239GPT-5 made this option configurable as one of `high`, `medium`, or `low`. With GPT-5.4, verbosity remains configurable and defaults to `medium`.
240
241When generating code with GPT-5.4, `medium` and `high` verbosity levels yield longer, more structured code with inline explanations, while `low` verbosity produces shorter, more concise code with minimal commentary.
242
243Control verbosity
244
245```bash
246curl --request POST \
247 --url https://api.openai.com/v1/responses \
248 --header "Authorization: Bearer $OPENAI_API_KEY" \
249 --header 'Content-type: application/json' \
250 --data '{
251 "model": "gpt-5.4",
252 "input": "What is the answer to the ultimate question of life, the universe, and everything?",
253 "text": {
254 "verbosity": "low"
255 }
256}'
257```
258
259```javascript
260import OpenAI from "openai";
261const openai = new OpenAI();
262
263const response = await openai.responses.create({
264 model: "gpt-5.4",
265 input: "What is the answer to the ultimate question of life, the universe, and everything?",
266 text: {
267 verbosity: "low"
268 }
269});
270
271console.log(response);
272```
273
274```python
275from openai import OpenAI
276client = OpenAI()
277
278response = client.responses.create(
279 model="gpt-5.4",
280 input="What is the answer to the ultimate question of life, the universe, and everything?",
281 text={
282 "verbosity": "low"
283 }
284)
285
286print(response)
287```
288
289
290You can still steer verbosity through prompting after setting it to `low` in the API. The verbosity parameter defines a general token range at the system prompt level, but the actual output is flexible to both developer and user prompts within that range.
291
292#### 1M context window
293
2941M token context window was introduced with GPT-5.4, making it easier to analyze entire codebases, long document collections, or extended agent trajectories in a single request.
295
296We have separate standard pricing for requests under 272K and over 272K tokens, available in the [pricing docs](https://developers.openai.com/api/docs/pricing). If you use [priority processing](https://developers.openai.com/api/docs/guides/priority-processing), any prompt above 272K tokens is automatically processed at standard rates.
297
298Long context pricing stacks with other pricing modifiers such as data residency and batch.
299
300We have different rate limits for requests under 272K tokens and over 272K tokens; this is available on the [GPT-5.4 model page](https://developers.openai.com/api/docs/models/gpt-5.4).
301
302## Using tools with GPT-5.4
303
304GPT-5.4 has been post-trained on specific tools. See the [tools docs](https://developers.openai.com/api/docs/guides/tools) for more specific guidance.
305
306### Computer use tool
307
308Computer use lets GPT-5.4 operate software through the user interface by inspecting screenshots and returning structured actions for your harness to execute. It is a good fit for browser or desktop workflows where a person could complete the task through the UI, such as navigating a site, filling out forms, or validating that a change actually worked.
309
310Use it in an isolated browser or VM, and keep a human in the loop for high-impact actions. The full guide covers the built-in Responses API loop, custom harness patterns, and code-execution-based setups.
311
312[
313
314<span slot="icon">
315 </span>
316 Learn how to run the built-in computer tool safely and integrate it with
317 your own harness.
318
319](https://developers.openai.com/api/docs/guides/tools-computer-use)
320
321### Tool search tool
322
323Tool search lets GPT-5.4 defer large tool surfaces until runtime so the model loads only the definitions it needs. This is most useful when you have many functions, namespaces, or MCP tools and want to reduce token usage, preserve cache performance, and improve latency without exposing every schema up front.
324
325Use hosted tool search when the candidate tools are already known at request time, or client-executed tool search when your application needs to decide what to load dynamically. The full guide also covers best practices for namespaces, MCP servers, and deferred loading.
326
327[
328
329<span slot="icon">
330 </span>
331 Learn how to defer tool definitions and load the right subset at runtime.
332
333](https://developers.openai.com/api/docs/guides/tools-tool-search)
334
335### Custom tools
336
337When the GPT-5 model family launched, we introduced a new capability called custom tools, which lets models send any raw text as tool call input but still constrain outputs if desired. This tool behavior remains true in GPT-5.4.
338 38
339[395. **Coding workflows need stronger orchestration:** GPT-5.5 is better suited to complex coding tasks that require planning, tool use, codebase navigation, verification, and multi-step execution. For coding agents, be explicit about reuse, subagent delegation, test expectations, acceptance criteria, and when to continue versus ask for help.
340 40
341<span slot="icon">41## Migration quickstart
342 </span>
343 Learn about custom tools in the function calling guide.
344 42
345](https://developers.openai.com/api/docs/guides/function-calling)43### Automated migration with Codex
346 44
347#### Freeform inputs45Codex can apply the recommended changes in this guide with the [OpenAI Docs Skill](https://github.com/openai/skills/tree/main/skills/.curated/openai-docs).
348
349Define your tool with `type: custom` to enable models to send plaintext inputs directly to your tools, rather than being limited to structured JSON. The model can send any raw text—code, SQL queries, shell commands, configuration files, or long-form prose—directly to your tool.
350
351```bash
352{
353 "type": "custom",
354 "name": "code_exec",
355 "description": "Executes arbitrary python code",
356}
357 46
47```text
48$openai-docs migrate this project to gpt-5.5
358```49```
359 50
360#### Constraining outputs51To use this skill in other coding agents, download it from the [OpenAI skills repository](https://github.com/openai/skills/tree/main/skills/.curated/openai-docs).
361
362GPT-5.4 supports context-free grammars (CFGs) for custom tools, letting you provide a Lark grammar to constrain outputs to a specific syntax or DSL. Attaching a CFG (e.g., a SQL or DSL grammar) ensures the assistant's text matches your grammar.
363
364This enables precise, constrained tool calls or structured responses and lets you enforce strict syntactic or domain-specific formats directly in GPT-5.4's function calling, improving control and reliability for complex or constrained domains.
365
366#### Best practices for custom tools
367
368- **Write concise, explicit tool descriptions**. The model chooses what to send based on your description; state clearly if you want it to always call the tool.
369- **Validate outputs on the server side**. Freeform strings are powerful but require safeguards against injection or unsafe commands.
370
371### Allowed tools
372
373The `allowed_tools` parameter under `tool_choice` lets you pass N tool definitions but restrict the model to only M (< N) of them. List your full toolkit in `tools`, and then use an `allowed_tools` block to name the subset and specify a mode—either `auto` (the model may pick any of those) or `required` (the model must invoke one).
374
375[
376
377<span slot="icon">
378 </span>
379 Learn about the allowed tools option in the function calling guide.
380
381](https://developers.openai.com/api/docs/guides/function-calling)
382
383By separating all possible tools from the subset that can be used _now_, you gain greater safety, predictability, and improved prompt caching. You also avoid brittle prompt engineering, such as hard-coded call order. GPT-5.4 dynamically invokes or requires specific functions mid-conversation while reducing the risk of unintended tool usage over long contexts.
384
385| | **Standard Tools** | **Allowed Tools** |
386| ---------------- | ----------------------------------------- | ------------------------------------------------------------- |
387| Model's universe | All tools listed under **`"tools": […]`** | Only the subset under **`"tools": […]`** in **`tool_choice`** |
388| Tool invocation | Model may or may not call any tool | Model restricted to (or required to call) chosen tools |
389| Purpose | Declare available capabilities | Constrain which capabilities are actually used |
390
391```bash
392 "tool_choice": {
393 "type": "allowed_tools",
394 "mode": "auto",
395 "tools": [
396 { "type": "function", "name": "get_weather" },
397 { "type": "function", "name": "search_docs" }
398 ]
399 }
400}'
401```
402
403For a more detailed overview of all of these new features, see the [prompt guidance for GPT-5.4](https://developers.openai.com/api/docs/guides/prompt-guidance).
404
405### Preambles
406
407Preambles are brief, user-visible explanations that GPT-5.4 generates before invoking any tool or function, outlining its intent or plan (e.g., “why I'm calling this tool”). They appear after the chain-of-thought and before the actual tool call, providing transparency into the model's reasoning and enhancing debuggability, user confidence, and fine-grained steerability.
408
409By letting GPT-5.4 “think out loud” before each tool call, preambles boost tool-calling accuracy (and overall task success) without bloating reasoning overhead. To enable preambles, add a system or developer instruction—for example: “Before you call a tool, explain why you are calling it.” GPT-5.4 prepends a concise rationale to each specified tool call. The model may also output multiple messages between tool calls, which can enhance the interaction experience—particularly for minimal reasoning or latency-sensitive use cases.
410
411For more on using preambles, see the [GPT-5 prompting cookbook](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_prompting_guide#tool-preambles).
412
413## Migration guidance
414
415GPT-5.4 is our best model yet, and it works best with the Responses API, which supports passing chain of thought (CoT) between turns to improve performance. Read below to migrate from your current model or API.
416
417### Migrating from other models to GPT-5.4
418
419Use the [OpenAI Docs
420 skill](https://github.com/openai/skills/tree/main/skills/.system/openai-docs)
421 when migrating existing prompts or workflows to GPT-5.4. It's available in our
422 public skills repository and the Codex desktop app.
423
424While the model should be close to a drop-in replacement for GPT-5.2, there are a few key changes to call out. See [Prompt guidance for GPT-5.4](https://developers.openai.com/api/docs/guides/prompt-guidance) for specific updates to make in your prompts.
425
426Using GPT-5 models with the Responses API provides improved intelligence because of the API's design. The Responses API can pass the previous turn's CoT to the model. This leads to fewer generated reasoning tokens, higher cache hit rates, and less latency. To learn more, see an [in-depth guide](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items) on the benefits of the Responses API.
427
428When migrating to GPT-5.4 from an older OpenAI model, start by experimenting with reasoning levels and prompting strategies. Based on our testing, we recommend using our [prompt optimizer](https://platform.openai.com/chat/edit?models=gpt-5.4&optimize=true)—which automatically updates your prompts for GPT-5.4 based on our best practices—and following this model-specific guidance:
429
430- **gpt-5.2**: `gpt-5.4` with default settings is meant to be a drop-in replacement.
431- **o3**: `gpt-5.4` with `medium` or `high` reasoning. Start with `medium` reasoning with prompt tuning, then increase to `high` if you aren't getting the results you want.
432- **gpt-4.1**: `gpt-5.4` with `none` reasoning. Start with `none` and tune your prompts; increase if you need better performance.
433- **o4-mini or gpt-4.1-mini**: `gpt-5.4-mini` with prompt tuning is a great replacement.
434- **gpt-4.1-nano**: `gpt-5.4-nano` with prompt tuning is a great replacement.
435
436### New `phase` parameter
437
438For long-running or tool-heavy GPT-5.4 flows in the Responses API, use the assistant message `phase` field to avoid early stopping and other misbehavior.
439
440`phase` is optional at the API level, but we highly recommend using it. Use `phase: "commentary"` for intermediate assistant updates (such as preambles before tool calls) and `phase: "final_answer"` for the completed answer. Do not add `phase` to user messages.
441
442If you use `previous_response_id`, that is usually the simplest path because
443 prior assistant state is preserved. If you replay assistant history manually,
444 preserve each original `phase` value.
445
446Missing or dropped `phase` can cause preambles to be treated as final answers
447in those workflows. For additional guidance and examples, see the [GPT-5.4
448prompting guide](https://developers.openai.com/api/docs/guides/prompt-guidance#phase-parameter).
449
450Round-trip assistant phase values
451
452```javascript
453import OpenAI from "openai";
454const client = new OpenAI();
455
456const response = await client.responses.create({
457 model: "gpt-5.4",
458 input: [
459 {
460 role: "assistant",
461 phase: "commentary",
462 content:
463 "I’ll inspect the logs and then summarize root cause and remediation.",
464 },
465 {
466 role: "assistant",
467 phase: "final_answer",
468 content: "Root cause: cache invalidation race.",
469 },
470 {
471 role: "user",
472 content: "Great—now give me a rollout-safe fix plan.",
473 },
474 ],
475});
476
477console.log(response.output_text);
478```
479
480```python
481from openai import OpenAI
482
483client = OpenAI()
484
485response = client.responses.create(
486 model="gpt-5.4",
487 input=[
488 {
489 "role": "assistant",
490 "phase": "commentary",
491 "content": "I’ll inspect the logs and then summarize root cause and remediation.",
492 },
493 {
494 "role": "assistant",
495 "phase": "final_answer",
496 "content": "Root cause: cache invalidation race.",
497 },
498 {
499 "role": "user",
500 "content": "Great—now give me a rollout-safe fix plan.",
501 },
502 ],
503)
504
505print(response.output_text)
506```
507
508
509### GPT-5.4 parameter compatibility
510
511The following parameters are **only supported** when using GPT-5.4 with reasoning effort set to `none`:
512
513- `temperature`
514- `top_p`
515- `logprobs`
516
517Requests to GPT-5.4 or GPT-5.2 with any other reasoning effort setting, or to older GPT-5 models (e.g., `gpt-5`, `gpt-5-mini`, `gpt-5-nano`) that include these fields will raise an error.
518
519To achieve similar results with reasoning effort set higher, or with another GPT-5 family model, try these alternative parameters:
520
521- **Reasoning depth:** `reasoning: { effort: "none" | "low" | "medium" | "high" | "xhigh" }`
522- **Output verbosity:** `text: { verbosity: "low" | "medium" | "high" }`
523- **Output length:** `max_output_tokens`
524
525### Migrating from Chat Completions to Responses API
526
527The biggest difference, and main reason to migrate from Chat Completions to the Responses API for GPT-5.4, is support for passing chain of thought (CoT) between turns. See a full [comparison of the APIs](https://developers.openai.com/api/docs/guides/responses-vs-chat-completions).
528
529Passing CoT exists only in the Responses API, and we've seen improved intelligence, fewer generated reasoning tokens, higher cache hit rates, and lower latency as a result of doing so. Most other parameters remain at parity, though the formatting is different. Here's how new parameters are handled differently between Chat Completions and the Responses API:
530
531**Reasoning effort**
532
533
534
535<div data-content-switcher-pane data-value="responses">
536 <div class="hidden">Responses API</div>
537 </div>
538 <div data-content-switcher-pane data-value="chat" hidden>
539 <div class="hidden">Chat Completions</div>
540 </div>
541
542
543
544**Verbosity**
545
546
547
548<div data-content-switcher-pane data-value="responses">
549 <div class="hidden">Responses API</div>
550 </div>
551 <div data-content-switcher-pane data-value="chat" hidden>
552 <div class="hidden">Chat Completions</div>
553 </div>
554
555
556
557**Custom tools**
558
559
560
561<div data-content-switcher-pane data-value="responses">
562 <div class="hidden">Responses API</div>
563 </div>
564 <div data-content-switcher-pane data-value="chat" hidden>
565 <div class="hidden">Chat Completions</div>
566 </div>
567
568
569
570## Prompting guidance
571
572We specifically designed GPT-5.4 to excel at coding and agentic tasks. We also recommend iterating on prompts for GPT-5.4 with the prompt optimizer.
573
574<div className="mt-4 flex flex-col gap-2">
575 [
576
577<span slot="icon">
578 </span>
579 Craft the perfect prompt for GPT-5.4 in the dashboard
580
581](https://platform.openai.com/chat/edit?optimize=true)
582
583[
584
585<span slot="icon">
586 </span>
587 Learn prompt patterns and migration tips for GPT-5.4
588
589](https://developers.openai.com/api/docs/guides/prompt-guidance)
590
591 <a href="https://cookbook.openai.com/examples/gpt-5/gpt-5_frontend">
592
593
594<span slot="icon">
595 </span>
596 See prompt samples specific to frontend development for GPT-5 family of
597 models
598
599
600 </a>
601</div>
602
603### GPT-5.4 is a reasoning model
604
605Reasoning models like GPT-5.4 break problems down step by step, producing an internal chain of thought that encodes their reasoning. To maximize performance, pass these reasoning items back to the model: this avoids re-reasoning and keeps interactions closer to the model's training distribution. In multi-turn conversations, passing a `previous_response_id` automatically makes earlier reasoning items available. This is especially important when using tools—for example, when a function call requires an extra round trip. In these cases, either include them with `previous_response_id` or add them directly to `input`.
606
607Learn more about reasoning models and how to get the most out of them in our [reasoning guide](https://developers.openai.com/api/docs/guides/reasoning).
608
609## Further reading
610
611[GPT-5.4 prompting guide](https://developers.openai.com/api/docs/guides/prompt-guidance)
612
613[GPT-5.3-Codex prompting guide](https://developers.openai.com/cookbook/examples/gpt-5/codex_prompting_guide)
614
615[GPT-5.4 blog post](https://openai.com/index/introducing-gpt-5-4/)
616
617[GPT-5 frontend guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_frontend)
618
619[GPT-5 model family: new features guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_new_params_and_tools)
620
621[Cookbook on reasoning models](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items)
622
623[Comparison of Responses API vs. Chat Completions](https://developers.openai.com/api/docs/guides/migrate-to-responses)
624
625## FAQ
626
6271. **How are these models integrated into ChatGPT?**
628
629 In ChatGPT, there are three models: GPT‑5 Instant, GPT‑5 Thinking, and GPT‑5 Pro. Based on the user's question, a routing layer selects the best model to use. Users can also invoke reasoning directly through the ChatGPT UI.
630 52
631 All three ChatGPT models (Instant, Thinking, and Pro) have a new knowledge cutoff of August 2025. For users, this means GPT-5.4 starts with a more current understanding of the world, so answers are more accurate and useful, with more relevant examples and context, even before turning to web search.53### API and model parameters
632 54
6331. **Will these models be supported in Codex?**55- Update the model slug to `gpt-5.5`.
56- Use the Responses API for any reasoning, tool-calling, or multi-turn use case.
57- Tune `reasoning.effort`. Use `low` for efficient reasoning, `medium` for a balanced point on the latency/performance curve, `high` for complex agentic tasks that require hard reasoning and where latency matters less, and `xhigh` for the hardest asynchronous agentic tasks or evals that test the bounds of model intelligence. See the [Reasoning models documentation](https://developers.openai.com/api/docs/guides/reasoning).
58- To configure for more concise responses, set `text.verbosity` to `low`. On GPT-5.5, this will result in proportionally more concise responses than `low` verbosity with GPT-5.4.
59- For tool-heavy or long-running workflows, verify that your application handles `phase`, preambles, and assistant-item replay correctly.
60- Benchmark against other models on accuracy, token consumption, and end-to-end latency.
634 61
635 Yes, `gpt-5.4` is the newest model that powers Codex and Codex CLI. You can also use this as a standalone model for building agentic coding applications.62### Prompting
636 63
6371. **How does GPT-5.4 compare to GPT-5.3-Codex?**64- State the expected outcome and success criteria.
65- Reduce or remove detailed step-by-step process guidance. Let GPT-5.5 choose the path unless the product requires that path.
66- Remove output schema definitions from the prompt where possible. Use [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) instead.
67- Optimize your prompt for caching: [static parts first, dynamic parts last](https://developers.openai.com/api/docs/guides/prompt-caching).
68- Drop the current date. The model is already aware of the current UTC date.
69- Review and optimize your prompts based on the guidance in [Prompting GPT-5.5](https://developers.openai.com/api/docs/guides/prompt-guidance?model=gpt-5.5).
638 70
639 [GPT-5.3-Codex](https://developers.openai.com/api/docs/models/gpt-5.3-codex) is specifically designed for use in coding environments such as Codex. GPT-5.4 is designed for both general-purpose work and coding, making it the better default when your workflow spans software engineering plus planning, writing, or other business tasks. GPT-5.3-Codex is only available in the Responses API and supports `low`, `medium`, `high`, and `xhigh` reasoning effort settings along with function calling, structured outputs, streaming, and prompt caching. It doesn't support all GPT-5.4 parameters or API surfaces.71## Using reasoning models
640 72
6411. **What is the deprecation plan for previous models?**73This guidance applies to GPT-5 series models and is worth revisiting whenever teams move workloads onto reasoning models. GPT-5.5 carries forward many capabilities that first appeared in earlier models, but they're still worth reviewing if you are moving from an earlier GPT-5 model, GPT-4.1, or a reasoning model such as o3.
642 74
643 Any model deprecations will be posted on our [deprecations page](https://developers.openai.com/api/docs/deprecations#page-top). We'll send advanced notice of any model deprecations.75Teams can overlook these features because they sit partly in API configuration and orchestration rather than in the prompt itself. Used together, the Responses API, reasoning controls, verbosity, structured outputs, prompt caching, tool design, hosted tools, and state management help reasoning models deliver their best intelligence, reliability, latency, and cost profile.
644 76
6451. **What are the reasoning efforts supported?**
646 - GPT 5 supports minimal, low, medium (default), and high.
647 - GPT 5.2 supports none (default), low, medium, and high.
648 - GPT 5.4 supports none (default), low, medium, high, and xhigh.
77- **Responses API:** GPT-5.5 works best in the [Responses API](https://developers.openai.com/api/docs/guides/migrate-to-responses). Use `previous_response_id` for multi-turn state handling. For stateless or Zero Data Retention flows, pass back the relevant returned output items each turn. See [Passing context from the previous response](https://developers.openai.com/api/docs/guides/conversation-state#passing-context-from-the-previous-response) for details.
78- **Reasoning effort:** Use `reasoning.effort` to choose between `low`, `medium`, `high`, or `xhigh`. The default is `medium`, but many workloads will perform well with `low`. Reserve `none` for use cases where low latency is more important than intelligence. See [Reasoning Models](https://developers.openai.com/api/docs/guides/reasoning) for detailed recommendations.
79- **Verbosity:** Use `text.verbosity` to control output length. Treat final answer length as separate from reasoning quality; specify word budgets, section counts, table widths, or JSON-only output where needed.
80- **Structured Outputs:** Avoid describing the expected output schema in the prompt. Use [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) for automatic validation and increased accuracy.
81- **Prompt caching:** [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) works automatically for eligible long prompts and can reduce latency and input-token cost. To maximize cache hits, keep stable content at the beginning of the request. Put dynamic user-specific context near the end. For repeated traffic with common prefixes, use `prompt_cache_key` consistently and track `usage.prompt_tokens_details.cached_tokens`.
82- **Tool calling:** GPT-5.5 supports the same tool-calling patterns as GPT-5.4, including function tools and tool-heavy agent workflows. Put most tool-specific guidance in the tool descriptions themselves: what the tool does, when to use it, required inputs, side effects, retry safety, and common error modes. Add tool-specific context to system instructions only when it applies across tools or materially changes the agent's operating policy.
83- **Hosted tools and tool search:** Prefer [OpenAI-hosted tools](https://developers.openai.com/api/docs/guides/tools) where they fit the workflow, such as web search, file search, code interpreter, image generation, and computer use. Hosted tools reduce custom orchestration burden and keep common tool patterns aligned with the Responses API and Agents SDK. Use custom function tools when you need to call your own systems, enforce domain-specific side effects, or expose internal business workflows. For large tool catalogs, consider using [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search) to defer tool definitions and load only the relevant subset.
84- **Tool preambles:** Preambles can improve chat UX because the user sees an initial, useful status update before the model generates the final response. They also make tool use easier to follow: the model can state what it's about to check or do, then continue from that same assistant state after tool results arrive.
85- **`phase` handling:** If your application manually manages Responses state by passing output items back each turn instead of using `previous_response_id`, preserve the `phase` parameter on returned assistant output items and pass it back unchanged. This is especially important when using reasoning effort, preambles, or repeated tool calls. See [Phase parameter](https://developers.openai.com/api/docs/guides/reasoning#phase-parameter).
86- **Compaction:** For long-running agents, use [conversation/state compaction](https://developers.openai.com/api/docs/guides/compaction) intentionally. Preserve completed actions, active assumptions, IDs, tool outcomes, unresolved blockers, and the next concrete goal.
87- **Agents SDK:** For new agentic systems, use the latest [Agents SDK](https://developers.openai.com/api/docs/guides/agents) patterns for tool orchestration, tracing, handoffs, and state management rather than rebuilding orchestration from scratch.
88- **Current date:** GPT-5.5 is aware of the current date in UTC. You don't need to add the current date to system instructions. Add explicit date or timezone context only when the application needs a business-specific timezone, policy-effective date, user-local date, or other non-UTC reference point.