guides/latest-model/gpt-5.4.md +0 −1383 deleted
File Deleted View Diff
1# Using GPT-5.4
2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4
5## Introduction
6
7[GPT-5.4](https://developers.openai.com/api/docs/models/gpt-5.4) was released as a frontier model for professional work across the API and Codex. It helps developers analyze complex information, build production software, and automate multi-step workflows.
8
9Within the GPT-5.4 generation, `gpt-5.4` is the general-purpose model for workflows that move between software engineering, reasoning, writing, and tool use.
10
11This guide covers key features of the GPT-5 model family and how to get the most out of GPT-5.4.
12
13## What's new
14
15Compared with the previous GPT-5.2 model, GPT-5.4 shows improvements in:
16
17- Coding, document understanding, tool use, and instruction following
18- Image perception and multimodal tasks
19- Long-running task execution and multi-step agent workflows
20- Token efficiency and end-to-end performance on tool-heavy workloads
21- Web search and multi-source synthesis for hard-to-locate information
22- Document-heavy and spreadsheet-heavy business workflows in customer service, analytics, and finance
23
24GPT-5.4 brings the coding capabilities of GPT-5.3-Codex to our flagship frontier model. Developers can generate production-quality code, build polished front-end UI, follow repo-specific patterns, and handle multi-file changes with fewer retries. It also has a strong out-of-the-box coding personality, so teams spend less time on prompt tuning.
25
26For agentic workloads, GPT-5.4 reduces end-to-end time across multi-step trajectories and often completes tasks with fewer tokens and tool calls. This makes agents more responsive and lowers the cost of operating complex workflows at scale in the API and Codex.
27
28### New features in GPT-5.4
29
30Like earlier GPT-5 models, GPT-5.4 supports custom tools, parameters to control verbosity and reasoning, and an allowed tools list. GPT-5.4 also introduces several capabilities that make it easier to build powerful agent systems, operate over larger bodies of information, and run more reliable automated workflows:
31
32- **`tool_search` in the API:** GPT-5.4 improves tool search for larger tool ecosystems by using deferred tool loading. This makes tools searchable, loads only the relevant definitions, reduces token usage, and improves tool selection accuracy in real deployments. Learn more in the [tool search guide](https://developers.openai.com/api/docs/guides/tools-tool-search).
33- **1M token context window:** GPT-5.4 supports up to a 1M token context window, making it easier to analyze entire codebases, long document collections, or extended agent trajectories in a single request. Read more in the [1M context window](#1m-context-window) section.
34- **Built-in computer use:** GPT-5.4 is the first mainline model with built-in computer-use capabilities, enabling agents to interact directly with software to complete, verify, and fix tasks in a build-run-verify-fix loop. Learn more in the [computer use guide](https://developers.openai.com/api/docs/guides/tools-computer-use).
35- **Native compaction support:** GPT-5.4 is the first mainline model trained to support compaction, enabling longer agent trajectories while preserving key context.
36
37## Model, API, and feature updates
38
39Within this model generation, `gpt-5.4` is the general-purpose model for both broad tasks and coding. For more difficult problems, `gpt-5.4-pro` uses more compute to think longer and provide more consistent answers.
40
41For smaller, faster variants, start with `gpt-5.4-mini` or `gpt-5.4-nano`.
42
43To help you pick the model that best fits your use case, consider these tradeoffs:
44
45| Variant | Best for |
46| ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
47| [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) | General-purpose work, including complex reasoning, broad world knowledge, and code-heavy or multi-step agentic tasks |
48| [`gpt-5.4-pro`](https://developers.openai.com/api/docs/models/gpt-5.4-pro) | Tough problems that may take longer to solve and need deeper reasoning |
49| [`gpt-5.4-mini`](https://developers.openai.com/api/docs/models/gpt-5.4-mini) | High-volume coding, computer use, and agent workflows that still need strong reasoning |
50| [`gpt-5.4-nano`](https://developers.openai.com/api/docs/models/gpt-5.4-nano) | High-throughput tasks where speed and cost matter most |
51
52### Lower reasoning effort
53
54The `reasoning.effort` parameter controls how many reasoning tokens the model generates before producing a response. Earlier reasoning models like o3 supported only `low`, `medium`, and `high`: `low` favored speed and fewer tokens, while `high` favored more thorough reasoning.
55
56GPT-5.2 and GPT-5.4 support `none` as their lowest reasoning effort for lower-latency interactions. It is the default setting for both models. If you need more thinking, slowly increase to `medium` and experiment with results.
57
58With reasoning effort set to `none`, prompting is important. To improve the model's reasoning quality, even with the default settings, encourage it to “think” or outline its steps before answering.
59
60Reasoning effort set to none
61
62```javascript
63import OpenAI from "openai";
64const openai = new OpenAI();
65
66const response = await openai.responses.create({
67 model: "gpt-5.4",
68 input:
69 "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
70 reasoning: {
71 effort: "none",
72 },
73});
74
75console.log(response);
76```
77
78```python
79from openai import OpenAI
80
81client = OpenAI()
82
83response = client.responses.create(
84 model="gpt-5.4",
85 input="Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
86 reasoning={"effort": "none"},
87)
88
89print(response)
90```
91
92```go
93package main
94
95import (
96 "context"
97 "fmt"
98
99 "github.com/openai/openai-go/v3"
100 "github.com/openai/openai-go/v3/responses"
101 "github.com/openai/openai-go/v3/shared"
102)
103
104func main() {
105 client := openai.NewClient()
106 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{
107 Model: "gpt-5.4",
108 Input: responses.ResponseNewParamsInputUnion{OfString: openai.String("Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?")},
109 Reasoning: shared.ReasoningParam{Effort: shared.ReasoningEffortNone},
110 })
111 if err != nil {
112 panic(err)
113 }
114 fmt.Println(response)
115}
116```
117
118```java
119import com.openai.client.OpenAIClient;
120import com.openai.client.okhttp.OpenAIOkHttpClient;
121import com.openai.models.Reasoning;
122import com.openai.models.ReasoningEffort;
123import com.openai.models.responses.ResponseCreateParams;
124
125ResponseCreateParams params =
126 ResponseCreateParams.builder()
127 .model("gpt-5.4")
128 .input("Explain the bug and propose a fix.")
129 .reasoning(Reasoning.builder().effort(ReasoningEffort.NONE).build())
130 .build();
131
132client.responses().create(params).output().stream()
133 .flatMap(item -> item.message().stream())
134 .flatMap(message -> message.content().stream())
135 .flatMap(content -> content.outputText().stream())
136 .forEach(text -> System.out.println(text.text()));
137```
138
139```csharp
140using OpenAI.Responses;
141#pragma warning disable OPENAI001
142
143string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;
144ResponsesClient client = new(key);
145
146CreateResponseOptions options = new()
147{
148 Model = "gpt-5.4",
149 ReasoningOptions = new ResponseReasoningOptions
150 {
151 ReasoningEffortLevel = ResponseReasoningEffortLevel.None,
152 },
153};
154options.InputItems.Add(
155 ResponseItem.CreateUserMessageItem(
156 "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?"
157 )
158);
159
160ResponseResult response = await client.CreateResponseAsync(options);
161Console.WriteLine(response.GetOutputText());
162```
163
164```ruby
165require "openai"
166
167client = OpenAI::Client.new
168response = client.responses.create(
169 model: "gpt-5.4",
170 reasoning: { effort: :minimal },
171 input: "Explain the bug and propose a fix."
172)
173puts(response.output_text)
174```
175
176```bash
177curl --request POST \
178 --url https://api.openai.com/v1/responses \
179 --header "Authorization: Bearer $OPENAI_API_KEY" \
180 --header 'Content-type: application/json' \
181 --data '{
182 "model": "gpt-5.4",
183 "input": "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
184 "reasoning": {
185 "effort": "none"
186 }
187}'
188```
189
190
191### Verbosity
192
193Verbosity determines how many output tokens are generated. Lowering the number of tokens reduces overall latency. While the model's reasoning approach stays mostly the same, the model finds ways to answer more concisely—which can either improve or diminish answer quality, depending on your use case. Here are some scenarios for both ends of the verbosity spectrum:
194
195- **High verbosity:** Use when you need the model to provide thorough explanations of documents or perform extensive code refactoring.
196- **Low verbosity:** Best for situations where you want concise answers or focused code generation, such as SQL queries.
197
198GPT-5 made this option configurable as one of `high`, `medium`, or `low`. With GPT-5.4, verbosity remains configurable and defaults to `medium`.
199
200When generating code with GPT-5.4, `medium` and `high` verbosity levels yield longer, more structured code with inline explanations, while `low` verbosity produces shorter, more concise code with minimal commentary.
201
202Control verbosity
203
204```javascript
205import OpenAI from "openai";
206const openai = new OpenAI();
207
208const response = await openai.responses.create({
209 model: "gpt-5.4",
210 input:
211 "What is the answer to the ultimate question of life, the universe, and everything?",
212 text: {
213 verbosity: "low",
214 },
215});
216
217console.log(response);
218```
219
220```python
221from openai import OpenAI
222
223client = OpenAI()
224
225response = client.responses.create(
226 model="gpt-5.4",
227 input="What is the answer to the ultimate question of life, the universe, and everything?",
228 text={"verbosity": "low"},
229)
230
231print(response)
232```
233
234```go
235package main
236
237import (
238 "context"
239 "fmt"
240
241 "github.com/openai/openai-go/v3"
242 "github.com/openai/openai-go/v3/responses"
243)
244
245func main() {
246 client := openai.NewClient()
247 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{
248 Model: "gpt-5.4",
249 Input: responses.ResponseNewParamsInputUnion{OfString: openai.String("What is the answer to the ultimate question of life, the universe, and everything?")},
250 Text: responses.ResponseTextConfigParam{Verbosity: responses.ResponseTextConfigVerbosityLow},
251 })
252 if err != nil {
253 panic(err)
254 }
255 fmt.Println(response)
256}
257```
258
259```java
260import com.openai.client.OpenAIClient;
261import com.openai.client.okhttp.OpenAIOkHttpClient;
262import com.openai.models.responses.ResponseCreateParams;
263import com.openai.models.responses.ResponseTextConfig;
264
265ResponseCreateParams params =
266 ResponseCreateParams.builder()
267 .model("gpt-5.4")
268 .input("Explain the bug and propose a fix.")
269 .text(ResponseTextConfig.builder().verbosity(ResponseTextConfig.Verbosity.LOW).build())
270 .build();
271
272client.responses().create(params).output().stream()
273 .flatMap(item -> item.message().stream())
274 .flatMap(message -> message.content().stream())
275 .flatMap(content -> content.outputText().stream())
276 .forEach(text -> System.out.println(text.text()));
277```
278
279```ruby
280require "openai"
281
282client = OpenAI::Client.new
283response = client.responses.create(
284 model: "gpt-5.4",
285 text: { verbosity: :low },
286 input: "Explain the bug and propose a fix."
287)
288puts(response.output_text)
289```
290
291```bash
292curl --request POST \
293 --url https://api.openai.com/v1/responses \
294 --header "Authorization: Bearer $OPENAI_API_KEY" \
295 --header 'Content-type: application/json' \
296 --data '{
297 "model": "gpt-5.4",
298 "input": "What is the answer to the ultimate question of life, the universe, and everything?",
299 "text": {
300 "verbosity": "low"
301 }
302}'
303```
304
305
306You can still steer verbosity through prompting after setting it to `low` in the API. The verbosity parameter defines a general token range at the system prompt level, but the actual output is flexible to both developer and user prompts within that range.
307
308#### 1M context window
309
3101M token context window was introduced with GPT-5.4, making it easier to analyze entire codebases, long document collections, or extended agent trajectories in a single request.
311
312We have separate standard pricing for requests under 272K and over 272K tokens, available in the [pricing docs](https://developers.openai.com/api/docs/pricing). If you use [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode), any prompt above 272K tokens is automatically processed at standard rates.
313
314Long context pricing stacks with other pricing modifiers such as data residency and batch.
315
316We have different rate limits for requests under 272K tokens and over 272K tokens; this is available on the [GPT-5.4 model page](https://developers.openai.com/api/docs/models/gpt-5.4).
317
318## Using tools with GPT-5.4
319
320GPT-5.4 has been post-trained on specific tools. See the [tools docs](https://developers.openai.com/api/docs/guides/tools) for more specific guidance.
321
322### Computer use tool
323
324Computer use lets GPT-5.4 operate software through the user interface by inspecting screenshots and returning structured actions for your harness to execute. It is a good fit for browser or desktop workflows where a person could complete the task through the UI, such as navigating a site, filling out forms, or validating that a change actually worked.
325
326Use it in an isolated browser or VM, and keep a human in the loop for high-impact actions. The full guide covers the built-in Responses API loop, custom harness patterns, and code-execution-based setups.
327
328[Computer use guide
329
330
331
332 Learn how to run the built-in computer tool safely and integrate it with
333 your own harness.](https://developers.openai.com/api/docs/guides/tools-computer-use)
334
335### Tool search tool
336
337Tool search lets GPT-5.4 defer large tool surfaces until runtime so the model loads only the definitions it needs. This is most useful when you have many functions, `namespaces`, or MCP tools and want to reduce token usage, preserve cache performance, and improve latency without exposing every schema up front.
338
339Use hosted tool search when the candidate tools are already known at request time, or client-executed tool search when your application needs to decide what to load dynamically. The full guide also covers best practices for `namespaces`, MCP servers, and deferred loading.
340
341[Tool search guide
342
343
344
345 Learn how to defer tool definitions and load the right subset at runtime.](https://developers.openai.com/api/docs/guides/tools-tool-search)
346
347### Custom tools
348
349When the GPT-5 model family launched, we introduced a new capability called custom tools, which lets models send any raw text as tool call input but still constrain outputs if desired. This tool behavior remains true in GPT-5.4.
350
351[Function calling guide
352
353
354
355 Learn about custom tools in the function calling guide.](https://developers.openai.com/api/docs/guides/function-calling)
356
357#### Freeform inputs
358
359Define your tool with `type: custom` to enable models to send plaintext inputs directly to your tools, rather than being limited to structured JSON. The model can send any raw text—code, SQL queries, shell commands, configuration files, or long-form prose—directly to your tool.
360
361```json
362{
363 "type": "custom",
364 "name": "code_exec",
365 "description": "Executes arbitrary python code"
366}
367```
368
369#### Constraining outputs
370
371GPT-5.4 supports context-free grammars (`CFGs`) for custom tools, letting you provide a Lark grammar to constrain outputs to a specific syntax or DSL. Attaching a CFG, for example a SQL or DSL grammar, ensures the assistant's text matches your grammar.
372
373This enables precise, constrained tool calls or structured responses and lets you enforce strict syntactic or domain-specific formats directly in GPT-5.4's function calling, improving control and reliability for complex or constrained domains.
374
375#### Best practices for custom tools
376
377- **Write concise, explicit tool descriptions.** The model chooses what to send based on your description; state explicitly if you want it to always call the tool.
378- **Validate outputs on the server side**. Freeform strings are powerful but require safeguards against injection or unsafe commands.
379
380### Allowed tools
381
382The `allowed_tools` parameter under `tool_choice` lets you pass N tool definitions but restrict the model to only M (< N) of them. List your full toolkit in `tools`, and then use an `allowed_tools` block to name the subset and specify a mode—either `auto` (the model may pick any of those) or `required` (the model must invoke one).
383
384[Function calling guide
385
386
387
388 Learn about the allowed tools option in the function calling guide.](https://developers.openai.com/api/docs/guides/function-calling)
389
390By separating all possible tools from the subset that can be used _now_, you gain greater safety, predictability, and improved prompt caching. You also avoid brittle prompt engineering, such as hard-coded call order. GPT-5.4 dynamically invokes or requires specific functions mid-conversation while reducing the risk of unintended tool usage over long contexts.
391
392| | **Standard Tools** | **Allowed Tools** |
393| ---------------- | ----------------------------------------- | ------------------------------------------------------------- |
394| Model's universe | All tools listed under **`"tools": […]`** | Only the subset under **`"tools": […]`** in **`tool_choice`** |
395| Tool invocation | Model may or may not call any tool | Model restricted to (or required to call) chosen tools |
396| Purpose | Declare available capabilities | Constrain which capabilities are actually used |
397
398```json
399{
400 "tool_choice": {
401 "type": "allowed_tools",
402 "mode": "auto",
403 "tools": [
404 { "type": "function", "name": "get_weather" },
405 { "type": "function", "name": "search_docs" }
406 ]
407 }
408}
409```
410
411For a more detailed overview of all of these new features, see the [prompt guidance for GPT-5.4](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.4#prompting-best-practices).
412
413### Preambles
414
415Preambles are brief, user-visible explanations that GPT-5.4 generates before invoking any tool or function, outlining its intent or plan—for example, “why I'm calling this tool.” They appear after the chain of thought and before the actual tool call, making the model's reasoning easier to understand and debug while supporting precise steering.
416
417By letting GPT-5.4 “think out loud” before each tool call, preambles boost tool-calling accuracy (and overall task success) without bloating reasoning overhead. To enable preambles, add a system or developer instruction—for example: “Before you call a tool, explain why you are calling it.” GPT-5.4 adds a concise rationale to each specified tool call. The model may also output multiple messages between tool calls, which can enhance the interaction experience—particularly for minimal reasoning or latency-sensitive use cases.
418
419For more on using preambles, see the [GPT-5 prompting cookbook](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_prompting_guide#tool-preambles).
420
421## Migration quickstart
422
423GPT-5.4 works best with the Responses API, which supports preserving reasoning context between turns to improve performance. Read below to migrate from your current model or API.
424
425### Migrating from other models to GPT-5.4
426
427Use the [OpenAI Docs
428 skill](https://github.com/openai/skills/tree/main/skills/.system/openai-docs)
429 when migrating existing prompts or workflows to GPT-5.4. It's available in our
430 public skills repository and the Codex desktop app.
431
432While the model should be close to a drop-in replacement for GPT-5.2, there are a few key changes to call out. See [Prompt guidance for GPT-5.4](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.4#prompting-best-practices) for specific updates to make in your prompts.
433
434Using GPT-5 models with the Responses API provides improved intelligence because of the API design. The Responses API can pass the previous turn's CoT to the model. This leads to fewer generated reasoning tokens, higher cache hit rates, and less latency. To learn more, see an [in-depth guide](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items) on the benefits of the Responses API.
435
436When migrating to GPT-5.4 from an older OpenAI model, start by experimenting with reasoning levels and prompting strategies. Use the [prompt optimizer](https://platform.openai.com/chat/edit?models=gpt-5.4&optimize=true) to update your prompts for GPT-5.4 based on current best practices, then follow this model-specific guidance:
437
438- **`gpt-5.2`**: `gpt-5.4` with default settings is meant to be a drop-in replacement.
439- **o3**: `gpt-5.4` with `medium` or `high` reasoning. Start with `medium` reasoning with prompt tuning, then increase to `high` if you aren't getting the results you want.
440- **`gpt-4.1`**: `gpt-5.4` with `none` reasoning. Start with `none` and tune your prompts; increase if you need better performance.
441- **`o4-mini` or `gpt-4.1-mini`**: `gpt-5.4-mini` with prompt tuning is a great replacement.
442- **`gpt-4.1-nano`**: `gpt-5.4-nano` with prompt tuning is a great replacement.
443
444### New `phase` parameter
445
446For long-running or tool-heavy GPT-5.4 flows in the Responses API, use the assistant message `phase` field to avoid early stopping and other misbehavior.
447
448`phase` is optional at the API level, but we highly recommend using it. Use `phase: "commentary"` for intermediate assistant updates (such as preambles before tool calls) and `phase: "final_answer"` for the completed answer. Do not add `phase` to user messages.
449
450If you use `previous_response_id`, that is usually the simplest path because
451 prior assistant state is preserved. If you replay assistant history manually,
452 preserve each original `phase` value.
453
454Missing or dropped `phase` can cause preambles to be treated as final answers
455in those workflows. For additional guidance and examples, see the [GPT-5.4
456prompting guide](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.4#phase-parameter).
457
458Round-trip assistant phase values
459
460```javascript
461import OpenAI from "openai";
462const client = new OpenAI();
463
464const response = await client.responses.create({
465 model: "gpt-5.4",
466 input: [
467 {
468 role: "assistant",
469 phase: "commentary",
470 content:
471 "I’ll inspect the logs and then summarize root cause and remediation.",
472 },
473 {
474 role: "assistant",
475 phase: "final_answer",
476 content: "Root cause: cache invalidation race.",
477 },
478 {
479 role: "user",
480 content: "Great—now give me a rollout-safe fix plan.",
481 },
482 ],
483});
484
485console.log(response.output_text);
486```
487
488```python
489from openai import OpenAI
490
491client = OpenAI()
492
493response = client.responses.create(
494 model="gpt-5.4",
495 input=[
496 {
497 "role": "assistant",
498 "phase": "commentary",
499 "content": "I’ll inspect the logs and then summarize root cause and remediation.",
500 },
501 {
502 "role": "assistant",
503 "phase": "final_answer",
504 "content": "Root cause: cache invalidation race.",
505 },
506 {
507 "role": "user",
508 "content": "Great—now give me a rollout-safe fix plan.",
509 },
510 ],
511)
512
513print(response.output_text)
514```
515
516```go
517package main
518
519import (
520 "context"
521 "fmt"
522
523 "github.com/openai/openai-go/v3"
524 "github.com/openai/openai-go/v3/responses"
525)
526
527func main() {
528 client := openai.NewClient()
529 commentary := responses.ResponseInputItemParamOfMessage(
530 "I’ll inspect the logs and then summarize root cause and remediation.",
531 responses.EasyInputMessageRoleAssistant,
532 )
533 commentary.OfMessage.Phase = responses.EasyInputMessagePhaseCommentary
534 finalAnswer := responses.ResponseInputItemParamOfMessage(
535 "Root cause: cache invalidation race.",
536 responses.EasyInputMessageRoleAssistant,
537 )
538 finalAnswer.OfMessage.Phase = responses.EasyInputMessagePhaseFinalAnswer
539
540 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{
541 Model: "gpt-5.4",
542 Input: responses.ResponseNewParamsInputUnion{OfInputItemList: responses.ResponseInputParam{
543 commentary,
544 finalAnswer,
545 responses.ResponseInputItemParamOfMessage("Great—now give me a rollout-safe fix plan.", responses.EasyInputMessageRoleUser),
546 }},
547 })
548 if err != nil {
549 panic(err)
550 }
551 fmt.Println(response.OutputText())
552}
553```
554
555```java
556import com.openai.client.OpenAIClient;
557import com.openai.client.okhttp.OpenAIOkHttpClient;
558import com.openai.models.Reasoning;
559import com.openai.models.ReasoningEffort;
560import com.openai.models.responses.ResponseCreateParams;
561
562ResponseCreateParams params =
563 ResponseCreateParams.builder()
564 .model("gpt-5.4")
565 .input("Explain the bug and propose a fix.")
566 .reasoning(Reasoning.builder().effort(ReasoningEffort.MEDIUM).build())
567 .build();
568
569client.responses().create(params).output().stream()
570 .flatMap(item -> item.message().stream())
571 .flatMap(message -> message.content().stream())
572 .flatMap(content -> content.outputText().stream())
573 .forEach(text -> System.out.println(text.text()));
574```
575
576```ruby
577require "openai"
578
579client = OpenAI::Client.new
580response = client.responses.create(
581 model: "gpt-5.4",
582 reasoning: { effort: :medium },
583 input: "Explain the bug and propose a fix."
584)
585puts(response.output_text)
586```
587
588
589### GPT-5.4 parameter compatibility
590
591The following parameters are **only supported** when using GPT-5.4 with reasoning effort set to `none`:
592
593- `temperature`
594- `top_p`
595- `logprobs`
596
597Requests that include these fields will raise an error for GPT-5.4 or GPT-5.2 with any other reasoning effort setting, or for older GPT-5 models such as `gpt-5`, `gpt-5-mini`, or `gpt-5-nano`.
598
599To achieve similar results with reasoning effort set higher, or with another GPT-5 family model, try these alternative parameters:
600
601- **Reasoning depth:** `reasoning: { effort: "none" | "low" | "medium" | "high" | "xhigh" }`
602- **Output verbosity:** `text: { verbosity: "low" | "medium" | "high" }`
603- **Output length:** `max_output_tokens`
604
605### Migrating from Chat Completions to Responses API
606
607The biggest difference, and main reason to migrate from Chat Completions to the Responses API for GPT-5.4, is support for passing chain of thought (CoT) between turns. See a full [comparison of the APIs](https://developers.openai.com/api/docs/guides/migrate-to-responses).
608
609Passing CoT exists only in the Responses API, and we've seen improved intelligence, fewer generated reasoning tokens, higher cache hit rates, and lower latency as a result of doing so. Most other parameters remain at parity, though the formatting is different. Here's how new parameters are handled differently between Chat Completions and the Responses API:
610
611**Reasoning effort**
612
613
614
615Responses API
616
617 Generate response with reasoning effort set to none
618
619```bash
620curl --request POST \
621 --url https://api.openai.com/v1/responses \
622 --header "Authorization: Bearer $OPENAI_API_KEY" \
623 --header "Content-type: application/json" \
624 --data '{
625 "model": "gpt-5.4",
626 "input": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?",
627 "reasoning": {
628 "effort": "none"
629 }
630}'
631```
632
633
634
635
636
637
638Chat Completions
639
640 Generate response with reasoning effort set to none
641
642```bash
643curl --request POST \
644 --url https://api.openai.com/v1/chat/completions \
645 --header "Authorization: Bearer $OPENAI_API_KEY" \
646 --header "Content-type: application/json" \
647 --data '{
648 "model": "gpt-5.4",
649 "messages": [
650 {
651 "role": "user",
652 "content": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?"
653 }
654 ],
655 "reasoning_effort": "none"
656}'
657```
658
659
660
661**Verbosity**
662
663
664
665Responses API
666
667 Control verbosity
668
669```bash
670curl --request POST \
671 --url https://api.openai.com/v1/responses \
672 --header "Authorization: Bearer $OPENAI_API_KEY" \
673 --header "Content-type: application/json" \
674 --data '{
675 "model": "gpt-5.4",
676 "input": "What is the answer to the ultimate question of life, the universe, and everything?",
677 "text": {
678 "verbosity": "low"
679 }
680}'
681```
682
683
684
685
686
687
688Chat Completions
689
690 Control verbosity
691
692```bash
693curl --request POST \
694 --url https://api.openai.com/v1/chat/completions \
695 --header "Authorization: Bearer $OPENAI_API_KEY" \
696 --header "Content-type: application/json" \
697 --data '{
698 "model": "gpt-5.4",
699 "messages": [
700 {
701 "role": "user",
702 "content": "What is the answer to the ultimate question of life, the universe, and everything?"
703 }
704 ],
705 "verbosity": "low"
706}'
707```
708
709
710
711**Custom tools**
712
713
714
715Responses API
716
717 Custom tool call
718
719```bash
720curl --request POST \
721 --url https://api.openai.com/v1/responses \
722 --header "Authorization: Bearer $OPENAI_API_KEY" \
723 --header "Content-type: application/json" \
724 --data '{
725 "model": "gpt-5.4",
726 "input": "Use the code_exec tool to calculate the area of a circle with radius equal to the number of r letters in blueberry",
727 "tools": [
728 {
729 "type": "custom",
730 "name": "code_exec",
731 "description": "Executes arbitrary Python code"
732 }
733 ]
734}'
735```
736
737
738
739
740
741
742Chat Completions
743
744 Custom tool call
745
746```bash
747curl --request POST \
748 --url https://api.openai.com/v1/chat/completions \
749 --header "Authorization: Bearer $OPENAI_API_KEY" \
750 --header "Content-type: application/json" \
751 --data '{
752 "model": "gpt-5.4",
753 "messages": [
754 {
755 "role": "user",
756 "content": "Use the code_exec tool to calculate the area of a circle with radius equal to the number of r letters in blueberry"
757 }
758 ],
759 "tools": [
760 {
761 "type": "custom",
762 "custom": {
763 "name": "code_exec",
764 "description": "Executes arbitrary Python code"
765 }
766 }
767 ]
768}'
769```
770
771
772
773
774## Prompting best practices
775
776When troubleshooting cases where GPT-5.4 treats an intermediate update as the
777 final answer, verify your integration preserves the assistant message `phase`
778 field correctly. See [Phase parameter](#phase-parameter) for details.
779
780### Understand GPT-5.4 behavior
781
782#### Where GPT-5.4 is strongest
783
784GPT-5.4 tends to work especially well in these areas:
785
786- Strong personality and tone adherence, with less drift over long answers
787- Agentic workflow robustness, with a stronger tendency to stick with multi-step work, retry, and complete agent loops end to end
788- Evidence-rich synthesis, especially in long-context or multi-tool workflows
789- Instruction adherence in modular, skill-based, and block-structured prompts when the contract is explicit
790- Long-context analysis across large, messy, or multi-document inputs
791- Batched or parallel tool calling while maintaining tool-call accuracy
792- Spreadsheet, finance, and Excel workflows that need instruction following, formatting fidelity, and stronger self-verification
793
794#### Where explicit prompting still helps
795
796Even with those strengths, GPT-5.4 benefits from more explicit guidance in a few recurring patterns:
797
798- Low-context tool routing early in a session, when tool selection can be less reliable
799- Dependency-aware workflows that need explicit prerequisite and downstream-step checks
800- Reasoning effort selection, where higher effort is not always better and the right choice depends on task shape, not intuition
801- Research tasks that require disciplined source collection and consistent citations
802- Irreversible or high-impact actions that require verification before execution
803- Terminal or coding-agent environments where tool boundaries must stay clear
804
805These patterns are observed defaults, not guarantees. Start with the smallest prompt that passes your evals, and add blocks only when they fix a measured failure mode.
806
807### Use core prompt patterns
808
809#### Keep outputs compact and structured
810
811To improve token efficiency with GPT-5.4, constrain verbosity and enforce structured output through clear output contracts. In practice, this acts as an additional control layer alongside the `verbosity` parameter in the Responses API, allowing you to guide both how much the model writes and how it structures the output.
812
813```xml
814<output_contract>
815- Return exactly the sections requested, in the requested order.
816- If the prompt defines a preamble, analysis block, or working section, do not treat it as extra output.
817- Apply length limits only to the section they are intended for.
818- If a format is required (JSON, Markdown, SQL, XML), output only that format.
819</output_contract>
820
821<verbosity_controls>
822- Prefer concise, information-dense writing.
823- Avoid repeating the user's request.
824- Keep progress updates brief.
825- Do not shorten the answer so aggressively that required evidence, reasoning, or completion checks are omitted.
826</verbosity_controls>
827```
828
829#### Set clear defaults for follow-through
830
831Users often change the task, format, or tone mid-conversation. To keep the assistant aligned, define clear rules for when to proceed, when to ask, and how newer instructions override earlier defaults.
832
833Use a default follow-through policy like this:
834
835```xml
836<default_follow_through_policy>
837- If the user’s intent is clear and the next step is reversible and low-risk, proceed without asking.
838- Ask permission only if the next step is:
839 (a) irreversible,
840 (b) has external side effects (for example sending, purchasing, deleting, or writing to production), or
841 (c) requires missing sensitive information or a choice that would materially change the outcome.
842- If proceeding, briefly state what you did and what remains optional.
843</default_follow_through_policy>
844```
845
846Make instruction priority explicit:
847
848```xml
849<instruction_priority>
850- User instructions override default style, tone, formatting, and initiative preferences.
851- Safety, honesty, privacy, and permission constraints do not yield.
852- If a newer user instruction conflicts with an earlier one, follow the newer instruction.
853- Preserve earlier instructions that do not conflict.
854</instruction_priority>
855```
856
857Higher-priority developer or system instructions remain binding.
858
859**Guidance:** When instructions change mid-conversation, make the update explicit, scoped, and local. State what changed, what still applies, and whether the change affects the next turn or the rest of the conversation.
860
861#### Handle mid-conversation instruction updates
862
863For mid-conversation updates, use explicit, scoped steering messages that state:
864
8651. Scope
8662. Override
8673. Carry forward
868
869```text
870<task_update>
871For the next response only:
872- Do not complete the task.
873- Only produce a plan.
874- Keep it to 5 bullets.
875
876All earlier instructions still apply unless they conflict with this update.
877</task_update>
878```
879
880If the task itself changes, say so directly:
881
882```text
883<task_update>
884The task has changed.
885Previous task: complete the workflow.
886Current task: review the workflow and identify risks only.
887
888Rules for this turn:
889- Do not execute actions.
890- Do not call destructive tools.
891- Return exactly:
892 1. Main risks
893 2. Missing information
894 3. Recommended next step
895</task_update>
896```
897
898#### Make tool use persistent when correctness depends on it
899
900Use explicit rules to keep tool use thorough, dependency-aware, and appropriately paced, especially in workflows where later actions rely on earlier retrieval or verification. A common failure mode is skipping prerequisites because the right end state seems obvious.
901
902GPT-5.4 can be less reliable at tool routing early in a session, when context is still thin. Prompt for prerequisites, dependency checks, and exact tool intent.
903
904```xml
905<tool_persistence_rules>
906- Use tools whenever they materially improve correctness, completeness, or grounding.
907- Do not stop early when another tool call is likely to materially improve correctness or completeness.
908- Keep calling tools until:
909 (1) the task is complete, and
910 (2) verification passes (see <verification_loop>).
911- If a tool returns empty or partial results, retry with a different strategy.
912</tool_persistence_rules>
913```
914
915This is especially important for workflows where the final action depends on earlier lookup or retrieval steps. One of the most common failure modes is skipping prerequisites because the intended end state seems obvious.
916
917```xml
918<dependency_checks>
919- Before taking an action, check whether prerequisite discovery, lookup, or memory retrieval steps are required.
920- Do not skip prerequisite steps just because the intended final action seems obvious.
921- If the task depends on the output of a prior step, resolve that dependency first.
922</dependency_checks>
923```
924
925Prompt for parallelism when the work is independent and wall-clock matters. Prompt for sequencing when dependencies, ambiguity, or irreversible actions matter more than speed.
926
927```xml
928<parallel_tool_calling>
929- When multiple retrieval or lookup steps are independent, prefer parallel tool calls to reduce wall-clock time.
930- Do not parallelize steps that have prerequisite dependencies or where one result determines the next action.
931- After parallel retrieval, pause to synthesize the results before making more calls.
932- Prefer selective parallelism: parallelize independent evidence gathering, not speculative or redundant tool use.
933</parallel_tool_calling>
934```
935
936#### Force completeness on long-horizon tasks
937
938For multi-step workflows, a common failure mode is incomplete execution: the model finishes after partial coverage, misses items in a batch, or treats empty or narrow retrieval as final. GPT-5.4 becomes more reliable when the prompt defines explicit completion rules and recovery behavior.
939
940Coverage can be achieved through sequential or parallel retrieval, but completion rules should remain explicit either way.
941
942```xml
943<completeness_contract>
944- Treat the task as incomplete until all requested items are covered or explicitly marked [blocked].
945- Keep an internal checklist of required deliverables.
946- For lists, batches, or paginated results:
947 - determine expected scope when possible,
948 - track processed items or pages,
949 - confirm coverage before finalizing.
950- If any item is blocked by missing data, mark it [blocked] and state exactly what is missing.
951</completeness_contract>
952```
953
954For workflows where empty, partial, or noisy retrieval is common:
955
956```xml
957<empty_result_recovery>
958If a lookup returns empty, partial, or suspiciously narrow results:
959- do not immediately conclude that no results exist,
960- try at least one or two fallback strategies,
961 such as:
962 - alternate query wording,
963 - broader filters,
964 - a prerequisite lookup,
965 - or an alternate source or tool,
966- Only then report that no results were found, along with what you tried.
967</empty_result_recovery>
968```
969
970#### Add a verification loop before high-impact actions
971
972Once the workflow appears complete, add a lightweight verification step before returning the answer or taking an irreversible action. This helps catch requirement misses, grounding issues, and format drift before commit.
973
974```xml
975<verification_loop>
976Before finalizing:
977- Check correctness: does the output satisfy every requirement?
978- Check grounding: are factual claims backed by the provided context or tool outputs?
979- Check formatting: does the output match the requested schema or style?
980- Check safety and irreversibility: if the next step has external side effects, ask permission first.
981</verification_loop>
982```
983
984```xml
985<missing_context_gating>
986- If required context is missing, do NOT guess.
987- Prefer the appropriate lookup tool when the missing context is retrievable; ask a minimal clarifying question only when it is not.
988- If you must proceed, label assumptions explicitly and choose a reversible action.
989</missing_context_gating>
990```
991
992For agents that actively take actions, add a short execution frame:
993
994```xml
995<action_safety>
996- Pre-flight: summarize the intended action and parameters in 1-2 lines.
997- Execute via tool.
998- Post-flight: confirm the outcome and any validation that was performed.
999</action_safety>
1000```
1001
1002### Handle specialized workflows
1003
1004#### Choose image detail explicitly for vision and computer use
1005
1006If your workflow depends on visual precision, specify the image `detail` level in the prompt or integration instead of relying on `auto`. Use `high` for standard high-fidelity image understanding. Use `original` for large, dense, or spatially sensitive images, especially [computer use, localization, OCR, and click-accuracy tasks](https://developers.openai.com/api/docs/guides/tools-computer-use) on `gpt-5.4` and future models. Use `low` only when speed and cost matter more than fine detail. For more details on image detail levels, see the [Images and Vision guide](https://developers.openai.com/api/docs/guides/images-vision).
1007
1008#### Lock research and citations to retrieved evidence
1009
1010When citation quality matters, make both the source boundary and the format requirement explicit. This helps reduce fabricated references, unsupported claims, and citation-format drift.
1011
1012```xml
1013<citation_rules>
1014- Only cite sources retrieved in the current workflow.
1015- Never fabricate citations, URLs, IDs, or quote spans.
1016- Use exactly the citation format required by the host application.
1017- Attach citations to the specific claims they support, not only at the end.
1018</citation_rules>
1019```
1020
1021```xml
1022<grounding_rules>
1023- Base claims only on provided context or tool outputs.
1024- If sources conflict, state the conflict explicitly and attribute each side.
1025- If the context is insufficient or irrelevant, narrow the answer or say you cannot support the claim.
1026- If a statement is an inference rather than a directly supported fact, label it as an inference.
1027</grounding_rules>
1028```
1029
1030If your application requires inline citations, require inline citations. If it requires footnotes, require footnotes. The key is to lock the format and prevent the model from improvising unsupported references.
1031
1032#### Research mode
1033
1034Push GPT-5.4 into a disciplined research mode. Use this pattern for research, review, and synthesis tasks. Do not force it onto short execution tasks or simple deterministic transforms.
1035
1036```xml
1037<research_mode>
1038- Do research in 3 passes:
1039 1) Plan: list 3-6 sub-questions to answer.
1040 2) Retrieve: search each sub-question and follow 1-2 second-order leads.
1041 3) Synthesize: resolve contradictions and write the final answer with citations.
1042- Stop only when more searching is unlikely to change the conclusion.
1043</research_mode>
1044```
1045
1046If your host environment uses a specific research tool or requires a submit step, combine this with the host's finalization contract.
1047
1048#### Clamp strict output formats
1049
1050For SQL, JSON, or other parse-sensitive outputs, tell GPT-5.4 to emit only the target format and check it before finishing.
1051
1052```text
1053<structured_output_contract>
1054- Output only the requested format.
1055- Do not add prose or markdown fences unless they were requested.
1056- Validate that parentheses and brackets are balanced.
1057- Do not invent tables or fields.
1058- If required schema information is missing, ask for it or return an explicit error object.
1059</structured_output_contract>
1060```
1061
1062If you are extracting document regions or OCR boxes, define the coordinate system and add a drift check:
1063
1064```text
1065<bbox_extraction_spec>
1066- Use the specified coordinate format exactly, such as [x1,y1,x2,y2] normalized to 0..1.
1067- For each box, include page, label, text snippet, and confidence.
1068- Add a vertical-drift sanity check so boxes stay aligned with the correct line of text.
1069- If the layout is dense, process page by page and do a second pass for missed items.
1070</bbox_extraction_spec>
1071```
1072
1073#### Keep tool boundaries explicit in coding and terminal agents
1074
1075In coding agents, GPT-5.4 works better when the rules for shell access and file editing are unambiguous. This is especially important when you expose tools like [Shell](https://developers.openai.com/api/docs/guides/tools-shell) or [Apply patch](https://developers.openai.com/api/docs/guides/tools-apply-patch).
1076
1077#### User updates
1078
1079GPT-5.4 does well with brief, outcome-based updates. Reuse the user-updates pattern from the 5.2 guide, but pair it with explicit completion and verification requirements.
1080
1081Recommended update spec:
1082
1083```xml
1084<user_updates_spec>
1085- Only update the user when starting a new major phase or when something changes the plan.
1086- Each update: 1 sentence on outcome + 1 sentence on next step.
1087- Do not narrate routine tool calls.
1088- Keep the user-facing status short; keep the work exhaustive.
1089</user_updates_spec>
1090```
1091
1092For coding agents, see the Prompting patterns for coding tasks section below for more specific guidance.
1093
1094#### Prompting patterns for coding tasks
1095
1096**Autonomy and persistence**
1097
1098GPT-5.4 is generally more thorough end to end than earlier mainline models on coding and tool-use tasks, so you often need less explicit "verify everything" prompting. Still, for high-stakes changes such as production, migrations, or security work, keep a lightweight verification clause.
1099
1100```xml
1101<autonomy_and_persistence>
1102Persist until the task is fully handled end-to-end within the current turn whenever feasible: do not stop at analysis or partial fixes; carry changes through implementation, verification, and a clear explanation of outcomes unless the user explicitly pauses or redirects you.
1103
1104Unless the user explicitly asks for a plan, asks a question about the code, is brainstorming potential solutions, or some other intent that makes it clear that code should not be written, assume the user wants you to make code changes or run tools to solve the user's problem. In these cases, it's bad to output your proposed solution in a message, you should go ahead and actually implement the change. If you encounter challenges or blockers, you should attempt to resolve them yourself.
1105</autonomy_and_persistence>
1106```
1107
1108**Intermediary updates**
1109
1110Keep updates sparse and high-signal. In coding tasks, prefer updates at key points.
1111
1112```xml
1113<user_updates_spec>
1114- Intermediary updates go to the `commentary` channel.
1115- User updates are short updates while you are working. They are not final answers.
1116- Use 1-2 sentence updates to communicate progress and new information while you work.
1117- Do not begin responses with conversational interjections or meta commentary. Avoid openers such as acknowledgements ("Done -", "Got it", or "Great question") or similar framing.
1118- Before exploring or doing substantial work, send a user update explaining your understanding of the request and your first step. Avoid commenting on the request or starting with phrases such as "Got it" or "Understood."
1119- Provide updates roughly every 30 seconds while working.
1120- When exploring, explain what context you are gathering and what you learned. Vary sentence structure so the updates do not become repetitive.
1121- When working for a while, keep updates informative and varied, but stay concise.
1122- When work is substantial, provide a longer plan after you have enough context. This is the only update that may be longer than 2 sentences and may contain formatting.
1123- Before file edits, explain what you are about to change.
1124- While thinking, keep the user informed of progress without narrating every tool call. Even if you are not taking actions, send frequent progress updates rather than going silent, especially if you are thinking for more than a short stretch.
1125- Keep the tone of progress updates consistent with the assistant's overall personality.
1126</user_updates_spec>
1127```
1128
1129**Formatting**
1130
1131GPT-5.4 often defaults to more structured formatting and may overuse bullet lists. If you want a clean final response, explicitly clamp list shape.
1132
1133```xml
1134Never use nested bullets. Keep lists flat (single level). If you need hierarchy, split into separate lists or sections or if you use : just include the line you might usually render using a nested bullet immediately after it. For numbered lists, only use the `1. 2. 3.` style markers (with a period), never `1)`.
1135```
1136
1137**Frontend tasks**
1138
1139Use this only when additional frontend guidance is useful.
1140
1141```xml
1142<frontend_tasks>
1143When doing frontend design tasks, avoid generic, overbuilt layouts.
1144
1145Use these hard rules:
1146- One composition: The first viewport must read as one composition, not a dashboard, unless it is a dashboard.
1147- Brand first: On branded pages, the brand or product name must be a hero-level signal, not just nav text or an eyebrow. No headline should overpower the brand.
1148- Brand test: If the first viewport could belong to another brand after removing the nav, the branding is too weak.
1149- Full-bleed hero only: On landing pages and promotional surfaces, the hero image should usually be a dominant edge-to-edge visual plane or background. Do not default to inset hero images, side-panel hero images, rounded media cards, tiled collages, or floating image blocks unless the existing design system clearly requires them.
1150- Hero budget: The first viewport should usually contain only the brand, one headline, one short supporting sentence, one CTA group, and one dominant image. Do not place stats, schedules, event listings, address blocks, promos, "this week" callouts, metadata rows, or secondary marketing content there.
1151- No hero overlays: Do not place detached labels, floating badges, promo stickers, info chips, or callout boxes on top of hero media.
1152- Cards: Default to no cards. Never use cards in the hero unless they are the container for a user interaction. If removing a border, shadow, background, or radius does not hurt interaction or understanding, it should not be a card.
1153- One job per section: Each section should have one purpose, one headline, and usually one short supporting sentence.
1154- Real visual anchor: Imagery should show the product, place, atmosphere, or context.
1155- Reduce clutter: Avoid pill clusters, stat strips, icon rows, boxed promos, schedule snippets, and competing text blocks.
1156- Use motion to create presence and hierarchy, not noise. Ship 2-3 intentional motions for visually led work, and prefer Framer Motion when it is available.
1157
1158Exception: If working within an existing website or design system, preserve the established patterns, structure, and visual language.
1159</frontend_tasks>
1160```
1161
1162```xml
1163<terminal_tool_hygiene>
1164- Only run shell commands via the terminal tool.
1165- Never "run" tool names as shell commands.
1166- If a patch or edit tool exists, use it directly; do not attempt it in bash.
1167- After changes, run a lightweight verification step such as ls, tests, or a build before declaring the task done.
1168</terminal_tool_hygiene>
1169```
1170
1171#### Document localization and OCR boxes
1172
1173For bbox tasks, be explicit about coordinate conventions and add drift tests.
1174
1175```xml
1176<bbox_extraction_spec>
1177- Use the specified coordinate format exactly (for example [x1,y1,x2,y2] normalized 0..1).
1178- For each bbox, include: page, label, text snippet, confidence.
1179- Add a vertical-drift sanity check:
1180 - ensure bboxes align with the line of text (not shifted up or down).
1181- If dense layout, process page by page and do a second pass for missed items.
1182</bbox_extraction_spec>
1183```
1184
1185#### Use runtime and API integration notes
1186
1187For long-running or tool-heavy agents, the runtime contract matters as much as the prompt contract.
1188
1189##### Phase parameter
1190
1191For GPT-5.4, `gpt-5.3-codex`, and later Responses models, the `phase` field can
1192help in the small number of long-running or tool-heavy flows where preambles or
1193other intermediate assistant updates are mistaken for the final answer.
1194
1195- `phase` is optional at the API level, but it is highly recommended. Best-effort inference may exist server-side, but explicit round-tripping of `phase` is strictly better.
1196- Use `phase` for long-running or tool-heavy agents that may emit commentary before tool calls or before a final answer.
1197- Preserve `phase` when replaying prior assistant items so the model can distinguish working commentary from the completed answer. This matters most in multi-step flows with preambles, tool-related updates, or multiple assistant messages in the same turn.
1198- Do not add `phase` to user messages.
1199- If you use `previous_response_id`, that is usually the simplest path, since OpenAI can often recover prior state without manually replaying assistant items.
1200- If you replay assistant history yourself, preserve the original `phase` values.
1201- Missing or dropped `phase` can cause preambles to be interpreted as final answers and degrade behavior on those multi-step tasks.
1202
1203#### Preserve behavior in long sessions
1204
1205Compaction unlocks significantly longer effective context windows, where user conversations can persist for many turns without hitting context limits or long-context performance degradation, and agents can perform very long trajectories that exceed a typical context window for long-running, complex tasks.
1206
1207If you are using [Compaction](https://developers.openai.com/api/docs/guides/compaction) in the Responses API, compact after major milestones, treat compacted items as opaque state, and keep prompts functionally identical after compaction. The endpoint is ZDR compatible and returns an `encrypted_content` item that you can pass into future requests. GPT-5.4 tends to remain more coherent and reliable over longer, multi-turn conversations with fewer breakdowns as sessions grow.
1208
1209For more guidance, see the [`/responses/compact` API reference](https://developers.openai.com/api/reference/resources/responses/methods/compact).
1210
1211#### Control personality for customer-facing workflows
1212
1213GPT-5.4 can be steered more effectively when you separate persistent personality from per-response writing controls. This is especially useful for customer-facing workflows such as emails, support replies, announcements, and blog-style content.
1214
1215- **Personality (persistent):** sets the default tone, verbosity, and decision style across the session.
1216- **Writing controls (per response):** define the channel, register, formatting, and length for a specific artifact.
1217- **Reminder:** personality should not override task-specific output requirements. If the user asks for JSON, return JSON.
1218
1219For natural, high-quality prose, the highest-leverage controls are:
1220
1221- Give the model a clear persona.
1222- Specify the channel and emotional register.
1223- Explicitly ban formatting when you want prose.
1224- Use hard length limits.
1225
1226```xml
1227<personality_and_writing_controls>
1228- Persona: <one sentence>
1229- Channel: <Slack | email | memo | PRD | blog>
1230- Emotional register: <direct/calm/energized/etc.> + "not <overdo this>"
1231- Formatting: <ban bullets/headers/markdown if you want prose>
1232- Length: <hard limit, e.g. <=150 words or 3-5 sentences>
1233- Default follow-through: if the request is clear and low-risk, proceed without asking permission.
1234</personality_and_writing_controls>
1235```
1236
1237For more personality patterns you can lift directly, see the [Prompt Personalities cookbook](https://developers.openai.com/cookbook/examples/gpt-5/prompt_personalities).
1238
1239**Professional memo mode**
1240
1241For memos, reviews, and other professional writing tasks, general writing instructions are often not enough. These workflows benefit from explicit guidance on specificity, domain conventions, synthesis, and calibrated certainty.
1242
1243```xml
1244<memo_mode>
1245- Write in a polished, professional memo style.
1246- Use exact names, dates, entities, and authorities when supported by the record.
1247- Follow domain-specific structure if one is requested.
1248- Prefer precise conclusions over generic hedging.
1249- When uncertainty is real, tie it to the exact missing fact or conflicting source.
1250- Synthesize across documents rather than summarizing each one independently.
1251</memo_mode>
1252```
1253
1254This mode is especially useful for legal, policy, research, and executive-facing writing, where the goal is not just fluency, but disciplined synthesis and clear conclusions.
1255
1256### Tune reasoning and migration
1257
1258#### Treat reasoning effort as a last-mile knob
1259
1260Reasoning effort is not one-size-fits-all. Treat it as a last-mile tuning knob, not the primary way to improve quality. In many cases, stronger prompts, clear output contracts, and lightweight verification loops recover much of the performance teams might otherwise seek through higher reasoning settings.
1261
1262Recommended defaults:
1263
1264- `none`: Best for fast, cost-sensitive, latency-sensitive tasks where the model does not need to think.
1265- `low`: Works well for latency-sensitive tasks where a small amount of thinking can produce a meaningful accuracy gain, especially with complex instructions.
1266- `medium` or `high`: Reserve for tasks that truly require stronger reasoning and can absorb the latency and cost tradeoff. Choose between them based on how much performance gain your task gets from additional reasoning.
1267- `xhigh`: Avoid as a default unless your evals show clear benefits. It is best suited for long, agentic, reasoning-heavy tasks where maximum intelligence matters more than speed or cost.
1268
1269In practice, most teams should default to the `none`, `low`, or `medium` range.
1270
1271Start with `none` for execution-heavy workloads such as workflow steps, field extraction, support triage, and short structured transforms.
1272
1273Start with `medium` or higher for research-heavy workloads such as long-context synthesis, multi-document review, conflict resolution, and strategy writing. With `medium` and a well-engineered prompt, you can squeeze out a lot of performance.
1274
1275For GPT-5.4 workloads, `none` can already perform well on action-selection and tool-discipline tasks. If your workload depends on nuanced interpretation, such as implicit requirements, ambiguity, or cancelled-tool-call recovery, start with `low` or `medium` instead.
1276
1277Before increasing reasoning effort, first add:
1278
1279- `<completeness_contract>`
1280- `<verification_loop>`
1281- `<tool_persistence_rules>`
1282
1283If the model still feels too literal or stops at the first plausible answer, add an initiative nudge before raising reasoning effort:
1284
1285```xml
1286<dig_deeper_nudge>
1287- Don’t stop at the first plausible answer.
1288- Look for second-order issues, edge cases, and missing constraints.
1289- If the task is safety or accuracy critical, perform at least one verification step.
1290</dig_deeper_nudge>
1291```
1292
1293#### Migrate prompts to GPT-5.4 one change at a time
1294
1295Use the same one-change-at-a-time discipline as the 5.2 guide: switch model first, pin `reasoning_effort`, run evals, then iterate.
1296
1297These starting points work well for many migrations:
1298
1299| Current setup | Suggested GPT-5.4 start | Notes |
1300| ------------------------- | ---------------------------------- | ------------------------------------------------------------------- |
1301| `gpt-5.2` | Match the current reasoning effort | Preserve the existing latency and quality profile first, then tune. |
1302| `gpt-5.3-codex` | Match the current reasoning effort | For coding workflows, keep the reasoning effort the same. |
1303| `gpt-4.1` or `gpt-4o` | `none` | Keep snappy behavior, and increase only if evals regress. |
1304| Research-heavy assistants | `medium` or `high` | Use explicit research multi-pass and citation gating. |
1305| Long-horizon agents | `medium` or `high` | Add tool persistence and completeness accounting. |
1306
1307#### Small-model guidance for `gpt-5.4-mini` and `gpt-5.4-nano`
1308
1309`gpt-5.4-mini` and `gpt-5.4-nano` are highly steerable, but they are less likely than larger models to infer missing steps, resolve ambiguity implicitly, or package outputs the way you intended unless you specify that behavior directly. In practice, prompts for smaller models are often a bit longer and more explicit.
1310
1311**How `gpt-5.4-mini` differs**
1312
1313- `gpt-5.4-mini` is more literal and makes fewer assumptions.
1314- It is strong when the task is clearly structured, but weaker on implicit workflows and ambiguity handling.
1315- By default, it may try to keep the conversation going with a follow-up question unless you suppress that behavior explicitly.
1316
1317**Prompting `gpt-5.4-mini`**
1318
1319- Put critical rules first.
1320- Specify the full execution order when tool use or side effects matter.
1321- Do not rely on "you MUST" alone. Use structural scaffolding such as numbered steps, decision rules, and explicit action definitions.
1322- Separate "do the action" from "report the action."
1323- Show the correct flow, not just the final format.
1324- Define ambiguity behavior explicitly: when to ask, abstain, or proceed.
1325- Specify packaging directly: answer length, whether to ask a follow-up question, citation style, and section order.
1326- Be careful with `output nothing else`. Prefer scoped instructions such as `after the final JSON, output nothing further`.
1327
1328**Prompting `gpt-5.4-nano`**
1329
1330- Use `gpt-5.4-nano` only for narrow, well-bounded tasks.
1331- Prefer closed outputs: labels, enums, short JSON, or fixed templates.
1332- Avoid multi-step orchestration unless the flow is extremely constrained.
1333- Route ambiguous or planning-heavy tasks to a stronger model instead of over-prompting `gpt-5.4-nano`.
1334
1335**Good default pattern**
1336
13371. Task
13382. Critical rule
13393. Exact step order
13404. Edge cases or clarification behavior
13415. Output format
13426. One correct example
1343
1344**Avoid**
1345
1346- Implied next steps
1347- Unspecified edge cases
1348- Schema-only prompts for tool workflows
1349- Generic instructions without structure
1350
1351#### Web search and deep research
1352
1353If you are migrating a research agent in particular, make these prompt updates before increasing reasoning effort:
1354
1355- Add `<research_mode>`
1356- Add `<citation_rules>`
1357- Add `<empty_result_recovery>`
1358- Increase `reasoning_effort` one notch only after prompt fixes.
1359
1360You can start from the 5.2 research block and then layer in citation gating and finalization contracts as needed.
1361
1362GPT-5.4 performs especially well when the task requires multi-step evidence gathering, long-context synthesis, and explicit prompt contracts. In practice, the highest-leverage prompt changes are choosing reasoning effort by task shape, defining exact output and citation formats, adding dependency-aware tool rules, and making completion criteria explicit. The model is often strong out of the box, but it is most reliable when prompts clearly specify how to search, how to verify, and what counts as done.
1363
1364### Next steps
1365
1366- Review [Model, API, and feature updates](#model-api-and-feature-updates) for model capabilities, parameters, and API compatibility details.
1367- Read [Prompt engineering](https://developers.openai.com/api/docs/guides/prompt-engineering) for broader prompting strategies that apply across model families.
1368- Read [Compaction](https://developers.openai.com/api/docs/guides/compaction) if you are building long-running GPT-5.4 sessions in the Responses API.
1369
1370
1371## Further reading
1372
1373[GPT-5.3-Codex prompting guide](https://developers.openai.com/cookbook/examples/gpt-5/codex_prompting_guide)
1374
1375[GPT-5.4 blog post](https://openai.com/index/introducing-gpt-5-4/)
1376
1377[GPT-5 frontend guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_frontend)
1378
1379[GPT-5 model family: new features guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_new_params_and_tools)
1380
1381[Cookbook on reasoning models](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items)
1382
1383[Comparison of Responses API vs. Chat Completions](https://developers.openai.com/api/docs/guides/migrate-to-responses)