SpyBara
Go Premium

Documentation 2026-09-16 20:58 UTC to 2026-09-17 10:04 UTC

12 files changed +160 −7,399. View all changes and history on the product overview
2026
Tue 29 22:57 Mon 28 22:57 Sat 26 23:59 Fri 25 23:58 Thu 24 23:58 Wed 23 23:58 Tue 22 23:57 Mon 21 23:00 Sat 19 23:00 Fri 18 22:59 Thu 17 10:04 Wed 16 20:58 Tue 15 22:59 Mon 14 22:58 Sun 13 15:02 Fri 11 20:00 Thu 10 18:01 Wed 9 23:59 Sat 5 17:01 Fri 4 23:59 Thu 3 23:00 Wed 2 22:59

guides/latest-model/gpt-4.1.md +0 −1601 deleted

File Deleted View Diff

1# Using GPT-4.1

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7The GPT-4.1 family of models represents a significant step forward from GPT-4o in capabilities across coding, instruction following, and long context. In this prompting guide, we collate a series of important prompting tips derived from extensive internal testing to help developers fully leverage the improved abilities of this new model family.

8 

9Many typical best practices still apply to GPT-4.1, such as providing context examples, making instructions as specific and clear as possible, and inducing planning via prompting to maximize model intelligence. However, we expect that getting the most out of this model will require some prompt migration. GPT-4.1 is trained to follow instructions more closely and more literally than its predecessors, which tended to more liberally infer intent from user and system prompts. This also means, however, that GPT-4.1 is highly steerable and responsive to well-specified prompts - if model behavior is different from what you expect, a single sentence firmly and unequivocally clarifying your desired behavior is almost always sufficient to steer the model on course.

10 

11Please read on for prompt examples you can use as a reference, and remember that while this guidance is widely applicable, no advice is one-size-fits-all. AI engineering is inherently an empirical discipline, and large language models are inherently nondeterministic; in addition to following this guide, we advise building informative evals and iterating often to ensure your prompt engineering changes are yielding benefits for your use case.

12 

13## What's new

14 

15- Closer and more literal instruction following than previous GPT models

16- Stronger coding and long-context behavior

17- Better API-native tool use when schemas are passed through the `tools` field

18- Prompt migration guidance for agentic workflows and diff generation

19 

20## Migration quickstart

21 

22- Update the model slug to `gpt-4.1`.

23- Use either the Responses API or Chat Completions API, depending on your integration.

24- Remove reasoning-specific parameters; GPT-4.1 is a non-reasoning model.

25- Pass tool schemas through the API `tools` field instead of injecting tool definitions into the prompt.

26- Review prompts for literal instruction following, add explicit persistence and tool-use rules where needed, and validate changes with evals.

27 

28## Model, API, and feature updates

29 

30- The GPT-4.1 family includes `gpt-4.1`, `gpt-4.1-mini`, and `gpt-4.1-nano`.

31- GPT-4.1 has a 1M-token context window and low latency without a reasoning step.

32- The family supports the Responses API and Chat Completions API.

33- GPT-4.1 and GPT-4.1 mini support supervised fine-tuning.

34- Supported tools include function calling, web search, file search, image generation, code interpreter, and remote MCP.

35 

36 

37## Prompting best practices

38 

39### 1. Agentic Workflows

40 

41GPT-4.1 is a great place to build agentic workflows. In model training we emphasized providing a diverse range of agentic problem-solving trajectories, and our agentic harness for the model achieves state-of-the-art performance for non-reasoning models on SWE-bench Verified, solving 55% of problems.

42 

43### System Prompt Reminders

44 

45In order to fully utilize the agentic capabilities of GPT-4.1, we recommend including three key types of reminders in all agent prompts. The following prompts are optimized specifically for the agentic coding workflow, but can be easily modified for general agentic use cases.

46 

471. Persistence: this ensures the model understands it is entering a multi-message turn, and prevents it from prematurely yielding control back to the user. Our example is the following:

48 

49```text

50You are an agent - please keep going until the user’s query is completely resolved, before ending your turn and yielding back to the user. Only terminate your turn when you are sure that the problem is solved.

51```

52 

532. Tool-calling: this encourages the model to make full use of its tools, and reduces its likelihood of hallucinating or guessing an answer. Our example is the following:

54 

55```text

56If you are not sure about file content or codebase structure pertaining to the user’s request, use your tools to read files and gather the relevant information: do NOT guess or make up an answer.

57```

58 

593. Planning \[optional\]: if desired, this ensures the model explicitly plans and reflects upon each tool call in text, instead of completing the task by chaining together a series of only tool calls. Our example is the following:

60 

61```text

62You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.

63```

64 

65GPT-4.1 is trained to respond very closely to both user instructions and system prompts in the agentic setting. The model adhered closely to these three simple instructions and increased our internal SWE-bench Verified score by close to 20% \- so we highly encourage starting any agent prompt with clear reminders covering the three categories listed above. As a whole, we find that these three instructions transform the model from a chatbot-like state into a much more “eager” agent, driving the interaction forward autonomously and independently.

66 

67### Tool Calls

68 

69Compared to previous models, GPT-4.1 has undergone more training on effectively utilizing tools passed as arguments in an OpenAI API request. We encourage developers to exclusively use the tools field to pass tools, rather than manually injecting tool descriptions into your prompt and writing a separate parser for tool calls, as some have reported doing in the past. This is the best way to minimize errors and ensure the model remains in distribution during tool-calling trajectories \- in our own experiments, we observed a 2% increase in SWE-bench Verified pass rate when using API-parsed tool descriptions versus manually injecting the schemas into the system prompt.

70 

71Developers should name tools clearly to indicate their purpose and add a clear, detailed description in the "description" field of the tool. Similarly, for each tool param, lean on good naming and descriptions to ensure appropriate usage. If your tool is particularly complicated and you'd like to provide examples of tool usage, we recommend that you create an `# Examples` section in your system prompt and place the examples there, rather than adding them into the "description" field, which should remain thorough but relatively concise. Providing examples can be helpful to indicate when to use tools, whether to include user text alongside tool calls, and what parameters are appropriate for different inputs. Remember that you can use “Generate Anything” in the [Prompt Playground](https://platform.openai.com/playground) to get a good starting point for your new tool definitions.

72 

73### Prompting-Induced Planning & Chain-of-Thought

74 

75As mentioned already, developers can optionally prompt agents built with GPT-4.1 to plan and reflect between tool calls, instead of silently calling tools in an unbroken sequence. GPT-4.1 is not a reasoning model \- meaning that it does not produce an internal chain of thought before answering \- but in the prompt, a developer can induce the model to produce an explicit, step-by-step plan by using any variant of the Planning prompt component shown above. This can be thought of as the model “thinking out loud.” In our experimentation with the SWE-bench Verified agentic task, inducing explicit planning increased the pass rate by 4%.

76 

77### Sample Prompt: SWE-bench Verified

78 

79Below, we share the agentic prompt that we used to achieve our highest score on SWE-bench Verified, which features detailed instructions about workflow and problem-solving strategy. This general pattern can be used for any agentic task.

80 

81```python

82from openai import OpenAI

83 

84client = OpenAI()

85 

86SYS_PROMPT_SWEBENCH = """

87You will be tasked to fix an issue from an open-source repository.

88 

89Your thinking should be thorough and so it's fine if it's very long. You can think step by step before and after each action you decide to take.

90 

91You MUST iterate and keep going until the problem is solved.

92 

93You already have everything you need to solve this problem in the /testbed folder, even without internet connection. I want you to fully solve this autonomously before coming back to me.

94 

95Only terminate your turn when you are sure that the problem is solved. Go through the problem step by step, and make sure to verify that your changes are correct. NEVER end your turn without having solved the problem, and when you say you are going to make a tool call, make sure you ACTUALLY make the tool call, instead of ending your turn.

96 

97THE PROBLEM CAN DEFINITELY BE SOLVED WITHOUT THE INTERNET.

98 

99Take your time and think through every step - remember to check your solution rigorously and watch out for boundary cases, especially with the changes you made. Your solution must be perfect. If not, continue working on it. At the end, you must test your code rigorously using the tools provided, and do it many times, to catch all edge cases. If it is not robust, iterate more and make it perfect. Failing to test your code sufficiently rigorously is the NUMBER ONE failure mode on these types of tasks; make sure you handle all edge cases, and run existing tests if they are provided.

100 

101You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.

102 

103# Workflow

104 

105## High-Level Problem Solving Strategy

106 

1071. Understand the problem deeply. Carefully read the issue and think critically about what is required.

1082. Investigate the codebase. Explore relevant files, search for key functions, and gather context.

1093. Develop a clear, step-by-step plan. Break down the fix into manageable, incremental steps.

1104. Implement the fix incrementally. Make small, testable code changes.

1115. Debug as needed. Use debugging techniques to isolate and resolve issues.

1126. Test frequently. Run tests after each change to verify correctness.

1137. Iterate until the root cause is fixed and all tests pass.

1148. Reflect and validate comprehensively. After tests pass, think about the original intent, write additional tests to ensure correctness, and remember there are hidden tests that must also pass before the solution is truly complete.

115 

116Refer to the detailed sections below for more information on each step.

117 

118## 1. Deeply Understand the Problem

119Carefully read the issue and think hard about a plan to solve it before coding.

120 

121## 2. Codebase Investigation

122- Explore relevant files and directories.

123- Search for key functions, classes, or variables related to the issue.

124- Read and understand relevant code snippets.

125- Identify the root cause of the problem.

126- Validate and update your understanding continuously as you gather more context.

127 

128## 3. Develop a Detailed Plan

129- Outline a specific, simple, and verifiable sequence of steps to fix the problem.

130- Break down the fix into small, incremental changes.

131 

132## 4. Making Code Changes

133- Before editing, always read the relevant file contents or section to ensure complete context.

134- If a patch is not applied correctly, attempt to reapply it.

135- Make small, testable, incremental changes that logically follow from your investigation and plan.

136 

137## 5. Debugging

138- Make code changes only if you have high confidence they can solve the problem

139- When debugging, try to determine the root cause rather than addressing symptoms

140- Debug for as long as needed to identify the root cause and identify a fix

141- Use print statements, logs, or temporary code to inspect program state, including descriptive statements or error messages to understand what's happening

142- To test hypotheses, you can also add test statements or functions

143- Revisit your assumptions if unexpected behavior occurs.

144 

145## 6. Testing

146- Run tests frequently using `!python3 run_tests.py` (or equivalent).

147- After each change, verify correctness by running relevant tests.

148- If tests fail, analyze failures and revise your patch.

149- Write additional tests if needed to capture important behaviors or edge cases.

150- Ensure all tests pass before finalizing.

151 

152## 7. Final Verification

153- Confirm the root cause is fixed.

154- Review your solution for logic correctness and robustness.

155- Iterate until you are extremely confident the fix is complete and all tests pass.

156 

157## 8. Final Reflection and Additional Testing

158- Reflect carefully on the original intent of the user and the problem statement.

159- Think about potential edge cases or scenarios that may not be covered by existing tests.

160- Write additional tests that would need to pass to fully validate the correctness of your solution.

161- Run these new tests and ensure they all pass.

162- Be aware that there are additional hidden tests that must also pass for the solution to be successful.

163- Do not assume the task is complete just because the visible tests pass; continue refining until you are confident the fix is robust and comprehensive.

164"""

165 

166PYTHON_TOOL_DESCRIPTION = """This function is used to execute Python code or terminal commands in a stateful Jupyter notebook environment. python will respond with the output of the execution or time out after 60.0 seconds. Internet access for this session is disabled. Do not make external web requests or API calls as they will fail. Just as in a Jupyter notebook, you may also execute terminal commands by calling this function with a terminal command, prefaced with an exclamation mark.

167 

168In addition, for the purposes of this task, you can call this function with an `apply_patch` command as input. `apply_patch` effectively allows you to execute a diff/patch against a file, but the format of the diff specification is unique to this task, so pay careful attention to these instructions. To use the `apply_patch` command, you should pass a message of the following structure as "input":

169 

170%%bash

171apply_patch <<"EOF"

172*** Begin Patch

173[YOUR_PATCH]

174*** End Patch

175EOF

176 

177Where [YOUR_PATCH] is the actual content of your patch, specified in the following V4A diff format.

178 

179*** [ACTION] File: [path/to/file] -> ACTION can be one of Add, Update, or Delete.

180For each snippet of code that needs to be changed, repeat the following:

181[context_before] -> See below for further instructions on context.

182- [old_code] -> Precede the old code with a minus sign.

183+ [new_code] -> Precede the new, replacement code with a plus sign.

184[context_after] -> See below for further instructions on context.

185 

186For instructions on [context_before] and [context_after]:

187- By default, show 3 lines of code immediately above and 3 lines immediately below each change. If a change is within 3 lines of a previous change, do NOT duplicate the first change's [context_after] lines in the second change's [context_before] lines.

188- If 3 lines of context is insufficient to uniquely identify the snippet of code within the file, use the @@ operator to indicate the class or function to which the snippet belongs. For instance, we might have:

189@@ class BaseClass

190[3 lines of pre-context]

191- [old_code]

192+ [new_code]

193[3 lines of post-context]

194 

195- If a code block is repeated so many times in a class or function such that even a single @@ statement and 3 lines of context cannot uniquely identify the snippet of code, you can use multiple `@@` statements to jump to the right context. For instance:

196 

197@@ class BaseClass

198@@ def method():

199[3 lines of pre-context]

200- [old_code]

201+ [new_code]

202[3 lines of post-context]

203 

204Note, then, that we do not use line numbers in this diff format, as the context is enough to uniquely identify code. An example of a message that you might pass as "input" to this function, in order to apply a patch, is shown below.

205 

206%%bash

207apply_patch <<"EOF"

208*** Begin Patch

209*** Update File: pygorithm/searching/binary_search.py

210@@ class BaseClass

211@@ def search():

212- pass

213+ raise NotImplementedError()

214 

215@@ class Subclass

216@@ def search():

217- pass

218+ raise NotImplementedError()

219 

220*** End Patch

221EOF

222 

223File references can only be relative, NEVER ABSOLUTE. After the apply_patch command is run, Python will always say "Done!", regardless of whether the patch was successfully applied or not. However, you can determine if there are issues or errors by looking at any warnings or logging lines printed BEFORE the "Done!" is output.

224"""

225 

226python_bash_patch_tool = {

227 "type": "function",

228 "name": "python",

229 "description": PYTHON_TOOL_DESCRIPTION,

230 "parameters": {

231 "type": "object",

232 "strict": True,

233 "properties": {

234 "input": {

235 "type": "string",

236 "description": " The Python code, terminal command (prefaced by exclamation mark), or apply_patch command that you wish to execute.",

237 }

238 },

239 "required": ["input"],

240 },

241}

242 

243# Additional harness setup:

244# - Add your repo to /testbed

245# - Add your issue to the first user message

246# - Note: Even though we used a single tool for python, bash, and apply_patch, we generally recommend defining more granular tools that are focused on a single function

247 

248response = client.responses.create(

249 instructions=SYS_PROMPT_SWEBENCH,

250 model="gpt-4.1-2025-04-14",

251 tools=[python_bash_patch_tool],

252 input="Please answer the following question:\nBug: Typerror...",

253)

254 

255response.to_dict()["output"]

256```

257 

258```java

259import com.openai.client.OpenAIClient;

260import com.openai.client.okhttp.OpenAIOkHttpClient;

261import com.openai.core.JsonValue;

262import com.openai.models.responses.FunctionTool;

263import com.openai.models.responses.ResponseCreateParams;

264import java.util.List;

265import java.util.Map;

266 

267String agentInstructions =

268 """

269 You will be tasked to fix an issue from an open-source repository.

270 

271 Your thinking should be thorough and so it's fine if it's very long. You can think step by step before and after each action you decide to take.

272 

273 You MUST iterate and keep going until the problem is solved.

274 

275 You already have everything you need to solve this problem in the /testbed folder, even without internet connection. I want you to fully solve this autonomously before coming back to me.

276 

277 Only terminate your turn when you are sure that the problem is solved. Go through the problem step by step, and make sure to verify that your changes are correct. NEVER end your turn without having solved the problem, and when you say you are going to make a tool call, make sure you ACTUALLY make the tool call, instead of ending your turn.

278 

279 THE PROBLEM CAN DEFINITELY BE SOLVED WITHOUT THE INTERNET.

280 

281 Take your time and think through every step - remember to check your solution rigorously and watch out for boundary cases, especially with the changes you made. Your solution must be perfect. If not, continue working on it. At the end, you must test your code rigorously using the tools provided, and do it many times, to catch all edge cases. If it is not robust, iterate more and make it perfect. Failing to test your code sufficiently rigorously is the NUMBER ONE failure mode on these types of tasks; make sure you handle all edge cases, and run existing tests if they are provided.

282 

283 You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.

284 

285 # Workflow

286 

287 ## High-Level Problem Solving Strategy

288 

289 1. Understand the problem deeply. Carefully read the issue and think critically about what is required.

290 2. Investigate the codebase. Explore relevant files, search for key functions, and gather context.

291 3. Develop a clear, step-by-step plan. Break down the fix into manageable, incremental steps.

292 4. Implement the fix incrementally. Make small, testable code changes.

293 5. Debug as needed. Use debugging techniques to isolate and resolve issues.

294 6. Test frequently. Run tests after each change to verify correctness.

295 7. Iterate until the root cause is fixed and all tests pass.

296 8. Reflect and validate comprehensively. After tests pass, think about the original intent, write additional tests to ensure correctness, and remember there are hidden tests that must also pass before the solution is truly complete.

297 

298 Refer to the detailed sections below for more information on each step.

299 

300 ## 1. Deeply Understand the Problem

301 Carefully read the issue and think hard about a plan to solve it before coding.

302 

303 ## 2. Codebase Investigation

304 - Explore relevant files and directories.

305 - Search for key functions, classes, or variables related to the issue.

306 - Read and understand relevant code snippets.

307 - Identify the root cause of the problem.

308 - Validate and update your understanding continuously as you gather more context.

309 

310 ## 3. Develop a Detailed Plan

311 - Outline a specific, simple, and verifiable sequence of steps to fix the problem.

312 - Break down the fix into small, incremental changes.

313 

314 ## 4. Making Code Changes

315 - Before editing, always read the relevant file contents or section to ensure complete context.

316 - If a patch is not applied correctly, attempt to reapply it.

317 - Make small, testable, incremental changes that logically follow from your investigation and plan.

318 

319 ## 5. Debugging

320 - Make code changes only if you have high confidence they can solve the problem

321 - When debugging, try to determine the root cause rather than addressing symptoms

322 - Debug for as long as needed to identify the root cause and identify a fix

323 - Use print statements, logs, or temporary code to inspect program state, including descriptive statements or error messages to understand what's happening

324 - To test hypotheses, you can also add test statements or functions

325 - Revisit your assumptions if unexpected behavior occurs.

326 

327 ## 6. Testing

328 - Run tests frequently using `!python3 run_tests.py` (or equivalent).

329 - After each change, verify correctness by running relevant tests.

330 - If tests fail, analyze failures and revise your patch.

331 - Write additional tests if needed to capture important behaviors or edge cases.

332 - Ensure all tests pass before finalizing.

333 

334 ## 7. Final Verification

335 - Confirm the root cause is fixed.

336 - Review your solution for logic correctness and robustness.

337 - Iterate until you are extremely confident the fix is complete and all tests pass.

338 

339 ## 8. Final Reflection and Additional Testing

340 - Reflect carefully on the original intent of the user and the problem statement.

341 - Think about potential edge cases or scenarios that may not be covered by existing tests.

342 - Write additional tests that would need to pass to fully validate the correctness of your solution.

343 - Run these new tests and ensure they all pass.

344 - Be aware that there are additional hidden tests that must also pass for the solution to be successful.

345 - Do not assume the task is complete just because the visible tests pass; continue refining until you are confident the fix is robust and comprehensive.

346 """;

347String pythonToolDescription =

348 """

349 This function is used to execute Python code or terminal commands in a stateful Jupyter notebook environment. python will respond with the output of the execution or time out after 60.0 seconds. Internet access for this session is disabled. Do not make external web requests or API calls as they will fail. Just as in a Jupyter notebook, you may also execute terminal commands by calling this function with a terminal command, prefaced with an exclamation mark.

350 

351 In addition, for the purposes of this task, you can call this function with an `apply_patch` command as input. `apply_patch` effectively allows you to execute a diff/patch against a file, but the format of the diff specification is unique to this task, so pay careful attention to these instructions. To use the `apply_patch` command, you should pass a message of the following structure as "input":

352 

353 %%bash

354 apply_patch <<"EOF"

355 *** Begin Patch

356 [YOUR_PATCH]

357 *** End Patch

358 EOF

359 

360 Where [YOUR_PATCH] is the actual content of your patch, specified in the following V4A diff format.

361 

362 *** [ACTION] File: [path/to/file] -> ACTION can be one of Add, Update, or Delete.

363 For each snippet of code that needs to be changed, repeat the following:

364 [context_before] -> See below for further instructions on context.

365 - [old_code] -> Precede the old code with a minus sign.

366 + [new_code] -> Precede the new, replacement code with a plus sign.

367 [context_after] -> See below for further instructions on context.

368 

369 For instructions on [context_before] and [context_after]:

370 - By default, show 3 lines of code immediately above and 3 lines immediately below each change. If a change is within 3 lines of a previous change, do NOT duplicate the first change's [context_after] lines in the second change's [context_before] lines.

371 - If 3 lines of context is insufficient to uniquely identify the snippet of code within the file, use the @@ operator to indicate the class or function to which the snippet belongs. For instance, we might have:

372 @@ class BaseClass

373 [3 lines of pre-context]

374 - [old_code]

375 + [new_code]

376 [3 lines of post-context]

377 

378 - If a code block is repeated so many times in a class or function such that even a single @@ statement and 3 lines of context cannot uniquely identify the snippet of code, you can use multiple `@@` statements to jump to the right context. For instance:

379 

380 @@ class BaseClass

381 @@ def method():

382 [3 lines of pre-context]

383 - [old_code]

384 + [new_code]

385 [3 lines of post-context]

386 

387 Note, then, that we do not use line numbers in this diff format, as the context is enough to uniquely identify code. An example of a message that you might pass as "input" to this function, in order to apply a patch, is shown below.

388 

389 %%bash

390 apply_patch <<"EOF"

391 *** Begin Patch

392 *** Update File: pygorithm/searching/binary_search.py

393 @@ class BaseClass

394 @@ def search():

395 - pass

396 + raise NotImplementedError()

397 

398 @@ class Subclass

399 @@ def search():

400 - pass

401 + raise NotImplementedError()

402 

403 *** End Patch

404 EOF

405 

406 File references can only be relative, NEVER ABSOLUTE. After the apply_patch command is run, Python will always say "Done!", regardless of whether the patch was successfully applied or not. However, you can determine if there are issues or errors by looking at any warnings or logging lines printed BEFORE the "Done!" is output.

407 """;

408 

409ResponseCreateParams params =

410 ResponseCreateParams.builder()

411 .model("gpt-4.1-2025-04-14")

412 .instructions(agentInstructions)

413 .input("Please answer the following question:\nBug: Typerror...")

414 .addTool(

415 FunctionTool.builder()

416 .name("python")

417 .description(pythonToolDescription)

418 .parameters(

419 FunctionTool.Parameters.builder()

420 .putAdditionalProperty("type", JsonValue.from("object"))

421 .putAdditionalProperty(

422 "properties",

423 JsonValue.from(

424 Map.of(

425 "input",

426 Map.of(

427 "type", "string",

428 "description",

429 "The Python code, terminal command, or apply_patch command to execute."))))

430 .putAdditionalProperty("required", JsonValue.from(List.of("input")))

431 .putAdditionalProperty("additionalProperties", JsonValue.from(false))

432 .build())

433 .strict(true)

434 .build())

435 .build();

436 

437client.responses().create(params).output().forEach(System.out::println);

438```

439 

440```ruby

441require "openai"

442 

443client = OpenAI::Client.new

444response = client.responses.create(

445 model: "gpt-4.1",

446 instructions: "Act as a coding agent. Inspect the reported failure, identify the smallest correct change, and explain how you would verify it.",

447 input: "The parser rejects an empty optional field even though the schema permits it. Diagnose the likely validation bug."

448)

449 

450puts(response.output_text)

451```

452 

453 

454```text

455[{'id': 'msg_67fe92df26ac819182ffafce9ff4e4fc07c7e06242e51f8b',

456 'content': [{'annotations': [],

457 'text': "Thank you for the report, but “Typerror” is too vague for me to start debugging right away.\n\n**To make progress, I need to:**\n1. Find the exact error message text (e.g. `'TypeError: ...'`).\n2. Find which file and which line/function/class the error occurred in.\n3. Figure out what triggered the error (test file, usage, reproduction steps).\n4. Find the root cause and details.\n\n**Next steps:**\n- Investigate error/log/test output files for a Python `TypeError` message.\n- Examine the relevant code sections for problematic type usage.\n- If possible, reproduce the bug locally.\n\n**Plan:**\n- First, I will search for test files and log output in the `/testbed` directory that may contain the full error message and stack trace.\n\nLet’s start by listing the contents of the `/testbed` directory to look for clues.",

458 'type': 'output_text'}],

459 'role': 'assistant',

460 'status': 'completed',

461 'type': 'message'},

462 {'arguments': '{"input":"!ls -l /testbed"}',

463 'call_id': 'call_frnxyJgKi5TsBem0nR9Zuzdw',

464 'name': 'python',

465 'type': 'function_call',

466 'id': 'fc_67fe92e3da7081918fc18d5c96dddc1c07c7e06242e51f8b',

467 'status': 'completed'}]

468```

469 

470### 2. Long context

471 

472GPT-4.1 has a performant 1M token input context window, and is useful for a variety of long context tasks, including structured document parsing, re-ranking, selecting relevant information while ignoring irrelevant context, and performing multi-hop reasoning using context.

473 

474### Optimal Context Size

475 

476We observe very good performance on needle-in-a-haystack evaluations up to our full 1M token context, and we’ve observed very strong performance at complex tasks with a mix of both relevant and irrelevant code and other documents. However, long context performance can degrade as more items are required to be retrieved, or perform complex reasoning that requires knowledge of the state of the entire context (like performing a graph search, for example).

477 

478### Tuning Context Reliance

479 

480Consider the mix of external vs. internal world knowledge that might be required to answer your question. Sometimes it’s important for the model to use some of its own knowledge to connect concepts or make logical jumps, while in others it’s desirable to only use provided context

481 

482```text

483# Instructions

484// for internal knowledge

485- Only use the documents in the provided External Context to answer the User Query. If you don't know the answer based on this context, you must respond "I don't have the information needed to answer that", even if a user insists on you answering the question.

486// For internal and external knowledge

487- By default, use the provided external context to answer the User Query, but if other basic knowledge is needed to answer, and you're confident in the answer, you can use some of your own knowledge to help answer the question.

488```

489 

490### Prompt Organization

491 

492Especially in long context usage, placement of instructions and context can impact performance. If you have long context in your prompt, ideally place your instructions at both the beginning and end of the provided context, as we found this to perform better than only above or below. If you’d prefer to only have your instructions once, then above the provided context works better than below.

493 

494### 3. Chain of Thought

495 

496As mentioned above, GPT-4.1 is not a reasoning model, but prompting the model to think step by step (called “chain of thought”) can be an effective way for a model to break down problems into more manageable pieces, solve them, and improve overall output quality, with the tradeoff of higher cost and latency associated with using more output tokens. The model has been trained to perform well at agentic reasoning about and real-world problem solving, so it shouldn’t require much prompting to perform well.

497 

498We recommend starting with this basic chain-of-thought instruction at the end of your prompt:

499 

500```text

501...

502 

503First, think carefully step by step about what documents are needed to answer the query. Then, print out the TITLE and ID of each document. Then, format the IDs into a list.

504```

505 

506From there, you should improve your chain-of-thought (CoT) prompt by auditing failures in your particular examples and evals, and addressing systematic planning and reasoning errors with more explicit instructions. In the unconstrained CoT prompt, there may be variance in the strategies it tries, and if you observe an approach that works well, you can codify that strategy in your prompt. Generally speaking, errors tend to occur from misunderstanding user intent, insufficient context gathering or analysis, or insufficient or incorrect step by step thinking, so watch out for these and try to address them with more opinionated instructions.

507 

508Here is an example prompt instructing the model to focus more methodically on analyzing user intent and considering relevant context before proceeding to answer.

509 

510```text

511# Reasoning Strategy

5121. Query Analysis: Break down and analyze the query until you're confident about what it might be asking. Consider the provided context to help clarify any ambiguous or confusing information.

5132. Context Analysis: Carefully select and analyze a large set of potentially relevant documents. Optimize for recall - it's okay if some are irrelevant, but the correct documents must be in this list, otherwise your final answer will be wrong. Analysis steps for each:

514 a. Analysis: An analysis of how it may or may not be relevant to answering the query.

515 b. Relevance rating: [high, medium, low, none]

5163. Synthesis: summarize which documents are most relevant and why, including all documents with a relevance rating of medium or higher.

517 

518# User Question

519{user_question}

520 

521# External Context

522{external_context}

523 

524First, think carefully step by step about what documents are needed to answer the query, closely adhering to the provided Reasoning Strategy. Then, print out the TITLE and ID of each document. Then, format the IDs into a list.

525```

526 

527### 4. Instruction Following

528 

529GPT-4.1 exhibits outstanding instruction-following performance, which developers can leverage to precisely shape and control the outputs for their particular use cases. Developers often extensively prompt for agentic reasoning steps, response tone and voice, tool calling information, output formatting, topics to avoid, and more. However, since the model follows instructions more literally, developers may need to include explicit specification around what to do or not to do. Furthermore, existing prompts optimized for other models may not immediately work with this model, because existing instructions are followed more closely and implicit rules are no longer being as strongly inferred.

530 

531### Recommended Workflow

532 

533Here is our recommended workflow for developing and debugging instructions in prompts:

534 

5351. Start with an overall “Response Rules” or “Instructions” section with high-level guidance and bullet points.

5362. If you’d like to change a more specific behavior, add a section to specify more details for that category, like `# Sample Phrases`.

5373. If there are specific steps you’d like the model to follow in its workflow, add an ordered list and instruct the model to follow these steps.

5384. If behavior still isn’t working as expected:

539 1. Check for conflicting, underspecified, or wrong instructions and examples. If there are conflicting instructions, GPT-4.1 tends to follow the one closer to the end of the prompt.

540 2. Add examples that demonstrate desired behavior; ensure that any important behavior demonstrated in your examples are also cited in your rules.

541 3. It’s generally not necessary to use all-caps or other incentives like bribes or tips. We recommend starting without these, and only reaching for these if necessary for your particular prompt. Note that if your existing prompts include these techniques, it could cause GPT-4.1 to pay attention to it too strictly.

542 

543_Note that using your preferred AI-powered IDE can be very helpful for iterating on prompts, including checking for consistency or conflicts, adding examples, or making cohesive updates like adding an instruction and updating instructions to demonstrate that instruction._

544 

545### Common Failure Modes

546 

547These failure modes are not unique to GPT-4.1, but we share them here for general awareness and ease of debugging.

548 

549- Instructing a model to always follow a specific behavior can occasionally induce adverse effects. For instance, if told “you must call a tool before responding to the user,” models may hallucinate tool inputs or call the tool with null values if they do not have enough information. Adding “if you don’t have enough information to call the tool, ask the user for the information you need” should mitigate this.

550- When provided sample phrases, models can use those quotes verbatim and start to sound repetitive to users. Ensure you instruct the model to vary them as necessary.

551- Without specific instructions, some models can be eager to provide additional prose to explain their decisions, or output more formatting in responses than may be desired. Provide instructions and potentially examples to help mitigate.

552 

553### Example Prompt: Customer Service

554 

555This demonstrates best practices for a fictional customer service agent. Observe the diversity of rules, the specificity, the use of additional sections for greater detail, and an example to demonstrate precise behavior that incorporates all prior rules.

556 

557Try running the following notebook cell - you should see both a user message and tool call, and the user message should start with a greeting, then echo back their answer, then mention they're about to call a tool. Try changing the instructions to shape the model behavior, or trying other user messages, to test instruction following performance.

558 

559```python

560SYS_PROMPT_CUSTOMER_SERVICE = """You are a helpful customer service agent working for NewTelco, helping a user efficiently fulfill their request while adhering closely to provided guidelines.

561 

562# Instructions

563- Always greet the user with "Hi, you've reached NewTelco, how can I help you?"

564- Always call a tool before answering factual questions about the company, its offerings or products, or a user's account. Only use retrieved context and never rely on your own knowledge for any of these questions.

565 - However, if you don't have enough information to properly call the tool, ask the user for the information you need.

566- Escalate to a human if the user requests.

567- Do not discuss prohibited topics (politics, religion, controversial current events, medical, legal, or financial advice, personal conversations, internal company operations, or criticism of any people or company).

568- Rely on sample phrases whenever appropriate, but never repeat a sample phrase in the same conversation. Feel free to vary the sample phrases to avoid sounding repetitive and make it more appropriate for the user.

569- Always follow the provided output format for new messages, including citations for any factual statements from retrieved policy documents.

570- If you're going to call a tool, always message the user with an appropriate message before and after calling the tool.

571- Maintain a professional and concise tone in all responses, and use emojis between sentences.

572- If you've resolved the user's request, ask if there's anything else you can help with

573 

574# Precise Response Steps (for each response)

5751. If necessary, call tools to fulfill the user's desired action. Always message the user before and after calling a tool to keep them in the loop.

5762. In your response to the user

577 a. Use active listening and echo back what you heard the user ask for.

578 b. Respond appropriately given the above guidelines.

579 

580# Sample Phrases

581## Deflecting a Prohibited Topic

582- "I'm sorry, but I'm unable to discuss that topic. Is there something else I can help you with?"

583- "That's not something I'm able to provide information on, but I'm happy to help with any other questions you may have."

584 

585## Before calling a tool

586- "To help you with that, I'll just need to verify your information."

587- "Let me check that for you—one moment, please."

588- "I'll retrieve the latest details for you now."

589 

590## After calling a tool

591- "Okay, here's what I found: [response]"

592- "So here's what I found: [response]"

593 

594# Output Format

595- Always include your final response to the user.

596- When providing factual information from retrieved context, always include citations immediately after the relevant statement(s). Use the following citation format:

597 - For a single source: [NAME](ID)

598 - For multiple sources: [NAME](ID), [NAME](ID)

599- Only provide information about this company, its policies, its products, or the customer's account, and only if it is based on information provided in context. Do not answer questions outside this scope.

600 

601# Example

602## User

603Can you tell me about your family plan options?

604 

605## Assistant Response 1

606### Message

607"Hi, you've reached NewTelco, how can I help you? 😊🎉\n\nYou'd like to know about our family plan options. 🤝 Let me check that for you—one moment, please. 🚀"

608 

609### Tool Calls

610lookup_policy_document(topic="family plan options")

611 

612// After tool call, the assistant would follow up with:

613 

614## Assistant Response 2 (after tool call)

615### Message

616"Okay, here's what I found: 🎉 Our family plan allows up to 5 lines with shared data and a 10% discount for each additional line [Family Plan Policy](ID-010). 📱 Is there anything else I can help you with today? 😊"

617"""

618 

619get_policy_doc = {

620 "type": "function",

621 "name": "lookup_policy_document",

622 "description": "Tool to look up internal documents and policies by topic or keyword.",

623 "parameters": {

624 "strict": True,

625 "type": "object",

626 "properties": {

627 "topic": {

628 "type": "string",

629 "description": "The topic or keyword to search for in company policies or documents.",

630 },

631 },

632 "required": ["topic"],

633 "additionalProperties": False,

634 },

635}

636 

637get_user_acct = {

638 "type": "function",

639 "name": "get_user_account_info",

640 "description": "Tool to get user account information",

641 "parameters": {

642 "strict": True,

643 "type": "object",

644 "properties": {

645 "phone_number": {

646 "type": "string",

647 "description": "Formatted as '(xxx) xxx-xxxx'",

648 },

649 },

650 "required": ["phone_number"],

651 "additionalProperties": False,

652 },

653}

654 

655response = client.responses.create(

656 instructions=SYS_PROMPT_CUSTOMER_SERVICE,

657 model="gpt-4.1-2025-04-14",

658 tools=[get_policy_doc, get_user_acct],

659 input="How much will it cost for international service? I'm traveling to France.",

660 # input="Why was my last bill so high?"

661)

662 

663response.to_dict()["output"]

664```

665 

666```java

667import com.openai.client.OpenAIClient;

668import com.openai.client.okhttp.OpenAIOkHttpClient;

669import com.openai.core.JsonValue;

670import com.openai.models.responses.FunctionTool;

671import com.openai.models.responses.ResponseCreateParams;

672import java.util.List;

673import java.util.Map;

674 

675String customerServiceInstructions =

676 """

677 You are a helpful customer service agent working for NewTelco, helping a user efficiently fulfill their request while adhering closely to provided guidelines.

678 

679 # Instructions

680 - Always greet the user with "Hi, you've reached NewTelco, how can I help you?"

681 - Always call a tool before answering factual questions about the company, its offerings or products, or a user's account. Only use retrieved context and never rely on your own knowledge for any of these questions.

682 - However, if you don't have enough information to properly call the tool, ask the user for the information you need.

683 - Escalate to a human if the user requests.

684 - Do not discuss prohibited topics (politics, religion, controversial current events, medical, legal, or financial advice, personal conversations, internal company operations, or criticism of any people or company).

685 - Rely on sample phrases whenever appropriate, but never repeat a sample phrase in the same conversation. Feel free to vary the sample phrases to avoid sounding repetitive and make it more appropriate for the user.

686 - Always follow the provided output format for new messages, including citations for any factual statements from retrieved policy documents.

687 - If you're going to call a tool, always message the user with an appropriate message before and after calling the tool.

688 - Maintain a professional and concise tone in all responses, and use emojis between sentences.

689 - If you've resolved the user's request, ask if there's anything else you can help with

690 

691 # Precise Response Steps (for each response)

692 1. If necessary, call tools to fulfill the user's desired action. Always message the user before and after calling a tool to keep them in the loop.

693 2. In your response to the user

694 a. Use active listening and echo back what you heard the user ask for.

695 b. Respond appropriately given the above guidelines.

696 

697 # Sample Phrases

698 ## Deflecting a Prohibited Topic

699 - "I'm sorry, but I'm unable to discuss that topic. Is there something else I can help you with?"

700 - "That's not something I'm able to provide information on, but I'm happy to help with any other questions you may have."

701 

702 ## Before calling a tool

703 - "To help you with that, I'll just need to verify your information."

704 - "Let me check that for you—one moment, please."

705 - "I'll retrieve the latest details for you now."

706 

707 ## After calling a tool

708 - "Okay, here's what I found: [response]"

709 - "So here's what I found: [response]"

710 

711 # Output Format

712 - Always include your final response to the user.

713 - When providing factual information from retrieved context, always include citations immediately after the relevant statement(s). Use the following citation format:

714 - For a single source: [NAME](ID)

715 - For multiple sources: [NAME](ID), [NAME](ID)

716 - Only provide information about this company, its policies, its products, or the customer's account, and only if it is based on information provided in context. Do not answer questions outside this scope.

717 

718 # Example

719 ## User

720 Can you tell me about your family plan options?

721 

722 ## Assistant Response 1

723 ### Message

724 "Hi, you've reached NewTelco, how can I help you? 😊🎉

725 

726 You'd like to know about our family plan options. 🤝 Let me check that for you—one moment, please. 🚀"

727 

728 ### Tool Calls

729 lookup_policy_document(topic="family plan options")

730 

731 // After tool call, the assistant would follow up with:

732 

733 ## Assistant Response 2 (after tool call)

734 ### Message

735 "Okay, here's what I found: 🎉 Our family plan allows up to 5 lines with shared data and a 10% discount for each additional line [Family Plan Policy](ID-010). 📱 Is there anything else I can help you with today? 😊"

736 """;

737 

738ResponseCreateParams params =

739 ResponseCreateParams.builder()

740 .model("gpt-4.1-2025-04-14")

741 .instructions(customerServiceInstructions)

742 .input("How much will it cost for international service? I'm traveling to France.")

743 .addTool(

744 customerServiceTool(

745 "lookup_policy_document",

746 "Tool to look up internal documents and policies by topic or keyword.",

747 "topic",

748 "The topic or keyword to search for in company policies or documents."))

749 .addTool(

750 customerServiceTool(

751 "get_user_account_info",

752 "Tool to get user account information",

753 "phone_number",

754 "Formatted as '(xxx) xxx-xxxx'"))

755 .build();

756 

757client.responses().create(params).output().forEach(System.out::println);

758 

759private static FunctionTool customerServiceTool(

760 String name, String description, String parameter, String parameterDescription) {

761 return FunctionTool.builder()

762 .name(name)

763 .description(description)

764 .strict(true)

765 .parameters(

766 FunctionTool.Parameters.builder()

767 .putAdditionalProperty("type", JsonValue.from("object"))

768 .putAdditionalProperty(

769 "properties",

770 JsonValue.from(

771 Map.of(

772 parameter,

773 Map.of("type", "string", "description", parameterDescription))))

774 .putAdditionalProperty("required", JsonValue.from(List.of(parameter)))

775 .putAdditionalProperty("additionalProperties", JsonValue.from(false))

776 .build())

777 .build();

778}

779```

780 

781```ruby

782require "openai"

783 

784client = OpenAI::Client.new

785response = client.responses.create(

786 model: "gpt-4.1",

787 instructions: "You are a customer service assistant. Confirm the customer's goal, use only supplied account facts, and clearly explain the next action.",

788 input: "A customer says a replacement order still has not shipped. Draft a concise response."

789)

790 

791puts(response.output_text)

792```

793 

794 

795```text

796[{'id': 'msg_67fe92d431548191b7ca6cd604b4784b06efc5beb16b3c5e',

797 'content': [{'annotations': [],

798 'text': "Hi, you've reached NewTelco, how can I help you? 🌍✈️\n\nYou'd like to know the cost of international service while traveling to France. 🇫🇷 Let me check the latest details for you—one moment, please. 🕑",

799 'type': 'output_text'}],

800 'role': 'assistant',

801 'status': 'completed',

802 'type': 'message'},

803 {'arguments': '{"topic":"international service cost France"}',

804 'call_id': 'call_cF63DLeyhNhwfdyME3ZHd0yo',

805 'name': 'lookup_policy_document',

806 'type': 'function_call',

807 'id': 'fc_67fe92d5d6888191b6cd7cf57f707e4606efc5beb16b3c5e',

808 'status': 'completed'}]

809```

810 

811### 5. General Advice

812 

813### Prompt Structure

814 

815For reference, here is a good starting point for structuring your prompts.

816 

817```text

818# Role and Objective

819 

820# Instructions

821 

822## Sub-categories for more detailed instructions

823 

824# Reasoning Steps

825 

826# Output Format

827 

828# Examples

829## Example 1

830 

831# Context

832 

833# Final instructions and prompt to think step by step

834```

835 

836Add or remove sections to suit your needs, and experiment to determine what’s optimal for your usage.

837 

838### Delimiters

839 

840Here are some general guidelines for selecting the best delimiters for your prompt. Please refer to the Long Context section for special considerations for that context type.

841 

8421. Markdown: We recommend starting here, and using markdown titles for major sections and subsections (including deeper hierarchy, to H4+). Use inline backticks or backtick blocks to precisely wrap code, and standard numbered or bulleted lists as needed.

8432. XML: These also perform well, and we have improved adherence to information in XML with this model. XML is convenient to precisely wrap a section including start and end, add metadata to the tags for additional context, and enable nesting. Here is an example of using XML tags to nest examples in an example section, with inputs and outputs for each:

844 

845```text

846<examples>

847<example1 type="Abbreviate">

848<input>San Francisco</input>

849<output>- SF</output>

850</example1>

851</examples>

852```

853 

8543. JSON is highly structured and well understood by the model particularly in coding contexts. However it can be more verbose, and require character escaping that can add overhead.

855 

856Guidance specifically for adding a large number of documents or files to input context:

857 

858- XML performed well in our long context testing.

859 - Example: `<doc id='1' title='The Fox'>The quick brown fox jumps over the lazy dog</doc>`

860- This format, proposed by Lee et al. ([ref](https://arxiv.org/pdf/2406.13121)), also performed well in our long context testing.

861 - Example: `ID: 1 | TITLE: The Fox | CONTENT: The quick brown fox jumps over the lazy dog`

862- JSON performed particularly poorly.

863 - Example: `[{'id': 1, 'title': 'The Fox', 'content': 'The quick brown fox jumped over the lazy dog'}]`

864 

865The model is trained to robustly understand structure in a variety of formats. Generally, use your judgement and think about what will provide clear information and “stand out” to the model. For example, if you’re retrieving documents that contain lots of XML, an XML-based delimiter will likely be less effective.

866 

867### Caveats

868 

869- In some isolated cases we have observed the model being resistant to producing very long, repetitive outputs, for example, analyzing hundreds of items one by one. If this is necessary for your use case, instruct the model strongly to output this information in full, and consider breaking down the problem or using a more concise approach.

870- We have seen some rare instances of parallel tool calls being incorrect. We advise testing this, and considering setting the [parallel_tool_calls](https://developers.openai.com/api/reference/resources/responses/methods/create#responses-create-parallel_tool_calls) param to false if you’re seeing issues.

871 

872### Appendix: Generating and Applying File Diffs

873 

874Developers have provided us feedback that accurate and well-formed diff generation is a critical capability to power coding-related tasks. To this end, the GPT-4.1 family features substantially improved diff capabilities relative to previous GPT models. Moreover, while GPT-4.1 has strong performance generating diffs of any format given clear instructions and examples, we open-source here one recommended diff format, on which the model has been extensively trained. We hope that in particular for developers just starting out, that this will take much of the guesswork out of creating diffs yourself.

875 

876### Apply Patch

877 

878See the example below for a prompt that applies our recommended tool call correctly.

879 

880```python

881APPLY_PATCH_TOOL_DESC = """This is a custom utility that makes it more convenient to add, remove, move, or edit code files. `apply_patch` effectively allows you to execute a diff/patch against a file, but the format of the diff specification is unique to this task, so pay careful attention to these instructions. To use the `apply_patch` command, you should pass a message of the following structure as "input":

882 

883%%bash

884apply_patch <<"EOF"

885*** Begin Patch

886[YOUR_PATCH]

887*** End Patch

888EOF

889 

890Where [YOUR_PATCH] is the actual content of your patch, specified in the following V4A diff format.

891 

892*** [ACTION] File: [path/to/file] -> ACTION can be one of Add, Update, or Delete.

893For each snippet of code that needs to be changed, repeat the following:

894[context_before] -> See below for further instructions on context.

895- [old_code] -> Precede the old code with a minus sign.

896+ [new_code] -> Precede the new, replacement code with a plus sign.

897[context_after] -> See below for further instructions on context.

898 

899For instructions on [context_before] and [context_after]:

900- By default, show 3 lines of code immediately above and 3 lines immediately below each change. If a change is within 3 lines of a previous change, do NOT duplicate the first change’s [context_after] lines in the second change’s [context_before] lines.

901- If 3 lines of context is insufficient to uniquely identify the snippet of code within the file, use the @@ operator to indicate the class or function to which the snippet belongs. For instance, we might have:

902@@ class BaseClass

903[3 lines of pre-context]

904- [old_code]

905+ [new_code]

906[3 lines of post-context]

907 

908- If a code block is repeated so many times in a class or function such that even a single @@ statement and 3 lines of context cannot uniquely identify the snippet of code, you can use multiple `@@` statements to jump to the right context. For instance:

909 

910@@ class BaseClass

911@@ def method():

912[3 lines of pre-context]

913- [old_code]

914+ [new_code]

915[3 lines of post-context]

916 

917Note, then, that we do not use line numbers in this diff format, as the context is enough to uniquely identify code. An example of a message that you might pass as "input" to this function, in order to apply a patch, is shown below.

918 

919%%bash

920apply_patch <<"EOF"

921*** Begin Patch

922*** Update File: pygorithm/searching/binary_search.py

923@@ class BaseClass

924@@ def search():

925- pass

926+ raise NotImplementedError()

927 

928@@ class Subclass

929@@ def search():

930- pass

931+ raise NotImplementedError()

932 

933*** End Patch

934EOF

935"""

936 

937APPLY_PATCH_TOOL = {

938 "name": "apply_patch",

939 "description": APPLY_PATCH_TOOL_DESC,

940 "parameters": {

941 "type": "object",

942 "properties": {

943 "input": {

944 "type": "string",

945 "description": " The apply_patch command that you wish to execute.",

946 }

947 },

948 "required": ["input"],

949 },

950}

951```

952 

953```ruby

954require "json"

955 

956APPLY_PATCH_TOOL_DESC = <<~PROMPT

957 This is a custom utility that makes it more convenient to add, remove, move, or edit code files. `apply_patch` effectively allows you to execute a diff/patch against a file, but the format of the diff specification is unique to this task, so pay careful attention to these instructions. To use the `apply_patch` command, you should pass a message of the following structure as "input":

958 

959 %%bash

960 apply_patch <<"EOF"

961 *** Begin Patch

962 [YOUR_PATCH]

963 *** End Patch

964 EOF

965 

966 Where [YOUR_PATCH] is the actual content of your patch, specified in the following V4A diff format.

967 

968 *** [ACTION] File: [path/to/file] -> ACTION can be one of Add, Update, or Delete.

969 For each snippet of code that needs to be changed, repeat the following:

970 [context_before] -> See below for further instructions on context.

971 - [old_code] -> Precede the old code with a minus sign.

972 + [new_code] -> Precede the new, replacement code with a plus sign.

973 [context_after] -> See below for further instructions on context.

974 

975 For instructions on [context_before] and [context_after]:

976 - By default, show 3 lines of code immediately above and 3 lines immediately below each change. If a change is within 3 lines of a previous change, do NOT duplicate the first change’s [context_after] lines in the second change’s [context_before] lines.

977 - If 3 lines of context is insufficient to uniquely identify the snippet of code within the file, use the @@ operator to indicate the class or function to which the snippet belongs. For instance, we might have:

978 @@ class BaseClass

979 [3 lines of pre-context]

980 - [old_code]

981 + [new_code]

982 [3 lines of post-context]

983 

984 - If a code block is repeated so many times in a class or function such that even a single @@ statement and 3 lines of context cannot uniquely identify the snippet of code, you can use multiple `@@` statements to jump to the right context. For instance:

985 

986 @@ class BaseClass

987 @@ def method():

988 [3 lines of pre-context]

989 - [old_code]

990 + [new_code]

991 [3 lines of post-context]

992 

993 Note, then, that we do not use line numbers in this diff format, as the context is enough to uniquely identify code. An example of a message that you might pass as "input" to this function, in order to apply a patch, is shown below.

994 

995 %%bash

996 apply_patch <<"EOF"

997 *** Begin Patch

998 *** Update File: pygorithm/searching/binary_search.py

999 @@ class BaseClass

1000 @@ def search():

1001 - pass

1002 + raise NotImplementedError()

1003 

1004 @@ class Subclass

1005 @@ def search():

1006 - pass

1007 + raise NotImplementedError()

1008 

1009 *** End Patch

1010 EOF

1011 

1012PROMPT

1013 

1014tool = {

1015 name: "apply_patch",

1016 description: APPLY_PATCH_TOOL_DESC,

1017 parameters: {

1018 type: "object",

1019 properties: {

1020 input: {

1021 type: "string",

1022 description: "The apply_patch command to execute."

1023 }

1024 },

1025 required: ["input"]

1026 }

1027}

1028puts(JSON.generate(tool))

1029```

1030 

1031 

1032### Reference Implementation: apply_patch.py

1033 

1034Here’s a reference implementation of the apply_patch tool that we used as part of model training. You’ll need to make this an executable and available as \`apply_patch\` from the shell where the model will execute commands:

1035 

1036```python

1037#!/usr/bin/env python3

1038 

1039"""

1040A self-contained **pure-Python 3.9+** utility for applying human-readable

1041“pseudo-diff” patch files to a collection of text files.

1042"""

1043 

1044from __future__ import annotations

1045 

1046import pathlib

1047from collections.abc import Callable

1048from dataclasses import dataclass, field

1049from enum import Enum

1050 

1051 

1052# --------------------------------------------------------------------------- #

1053# Domain objects

1054# --------------------------------------------------------------------------- #

1055class ActionType(str, Enum):

1056 ADD = "add"

1057 DELETE = "delete"

1058 UPDATE = "update"

1059 

1060 

1061@dataclass

1062class FileChange:

1063 type: ActionType

1064 old_content: str | None = None

1065 new_content: str | None = None

1066 move_path: str | None = None

1067 

1068 

1069@dataclass

1070class Commit:

1071 changes: dict[str, FileChange] = field(default_factory=dict)

1072 

1073 

1074# --------------------------------------------------------------------------- #

1075# Exceptions

1076# --------------------------------------------------------------------------- #

1077class DiffError(ValueError):

1078 """Any problem detected while parsing or applying a patch."""

1079 

1080 

1081# --------------------------------------------------------------------------- #

1082# Helper dataclasses used while parsing patches

1083# --------------------------------------------------------------------------- #

1084@dataclass

1085class Chunk:

1086 orig_index: int = -1

1087 del_lines: list[str] = field(default_factory=list)

1088 ins_lines: list[str] = field(default_factory=list)

1089 

1090 

1091@dataclass

1092class PatchAction:

1093 type: ActionType

1094 new_file: str | None = None

1095 chunks: list[Chunk] = field(default_factory=list)

1096 move_path: str | None = None

1097 

1098 

1099@dataclass

1100class Patch:

1101 actions: dict[str, PatchAction] = field(default_factory=dict)

1102 

1103 

1104# --------------------------------------------------------------------------- #

1105# Patch text parser

1106# --------------------------------------------------------------------------- #

1107@dataclass

1108class Parser:

1109 current_files: dict[str, str]

1110 lines: list[str]

1111 index: int = 0

1112 patch: Patch = field(default_factory=Patch)

1113 fuzz: int = 0

1114 

1115 # ------------- low-level helpers -------------------------------------- #

1116 def _cur_line(self) -> str:

1117 if self.index >= len(self.lines):

1118 raise DiffError("Unexpected end of input while parsing patch")

1119 return self.lines[self.index]

1120 

1121 @staticmethod

1122 def _norm(line: str) -> str:

1123 """Strip CR so comparisons work for both LF and CRLF input."""

1124 return line.rstrip("\r")

1125 

1126 # ------------- scanning convenience ----------------------------------- #

1127 def is_done(self, prefixes: tuple[str, ...] | None = None) -> bool:

1128 if self.index >= len(self.lines):

1129 return True

1130 if (

1131 prefixes

1132 and len(prefixes) > 0

1133 and self._norm(self._cur_line()).startswith(prefixes)

1134 ):

1135 return True

1136 return False

1137 

1138 def startswith(self, prefix: str | tuple[str, ...]) -> bool:

1139 return self._norm(self._cur_line()).startswith(prefix)

1140 

1141 def read_str(self, prefix: str) -> str:

1142 """

1143 Consume the current line if it starts with *prefix* and return the text

1144 **after** the prefix. Raises if prefix is empty.

1145 """

1146 if prefix == "":

1147 raise ValueError("read_str() requires a non-empty prefix")

1148 if self._norm(self._cur_line()).startswith(prefix):

1149 text = self._cur_line()[len(prefix) :]

1150 self.index += 1

1151 return text

1152 return ""

1153 

1154 def read_line(self) -> str:

1155 """Return the current raw line and advance."""

1156 line = self._cur_line()

1157 self.index += 1

1158 return line

1159 

1160 # ------------- public entry point -------------------------------------- #

1161 def parse(self) -> None:

1162 while not self.is_done(("*** End Patch",)):

1163 # ---------- UPDATE ---------- #

1164 path = self.read_str("*** Update File: ")

1165 if path:

1166 if path in self.patch.actions:

1167 raise DiffError(f"Duplicate update for file: {path}")

1168 move_to = self.read_str("*** Move to: ")

1169 if path not in self.current_files:

1170 raise DiffError(f"Update File Error - missing file: {path}")

1171 text = self.current_files[path]

1172 action = self._parse_update_file(text)

1173 action.move_path = move_to or None

1174 self.patch.actions[path] = action

1175 continue

1176 

1177 # ---------- DELETE ---------- #

1178 path = self.read_str("*** Delete File: ")

1179 if path:

1180 if path in self.patch.actions:

1181 raise DiffError(f"Duplicate delete for file: {path}")

1182 if path not in self.current_files:

1183 raise DiffError(f"Delete File Error - missing file: {path}")

1184 self.patch.actions[path] = PatchAction(type=ActionType.DELETE)

1185 continue

1186 

1187 # ---------- ADD ---------- #

1188 path = self.read_str("*** Add File: ")

1189 if path:

1190 if path in self.patch.actions:

1191 raise DiffError(f"Duplicate add for file: {path}")

1192 if path in self.current_files:

1193 raise DiffError(f"Add File Error - file already exists: {path}")

1194 self.patch.actions[path] = self._parse_add_file()

1195 continue

1196 

1197 raise DiffError(f"Unknown line while parsing: {self._cur_line()}")

1198 

1199 if not self.startswith("*** End Patch"):

1200 raise DiffError("Missing *** End Patch sentinel")

1201 self.index += 1 # consume sentinel

1202 

1203 # ------------- section parsers ---------------------------------------- #

1204 def _parse_update_file(self, text: str) -> PatchAction:

1205 action = PatchAction(type=ActionType.UPDATE)

1206 lines = text.split("\n")

1207 index = 0

1208 while not self.is_done(

1209 (

1210 "*** End Patch",

1211 "*** Update File:",

1212 "*** Delete File:",

1213 "*** Add File:",

1214 "*** End of File",

1215 )

1216 ):

1217 def_str = self.read_str("@@ ")

1218 section_str = ""

1219 if not def_str and self._norm(self._cur_line()) == "@@":

1220 section_str = self.read_line()

1221 

1222 if not (def_str or section_str or index == 0):

1223 raise DiffError(f"Invalid line in update section:\n{self._cur_line()}")

1224 

1225 if def_str.strip():

1226 found = False

1227 if def_str not in lines[:index]:

1228 for i, s in enumerate(lines[index:], index):

1229 if s == def_str:

1230 index = i + 1

1231 found = True

1232 break

1233 if not found and def_str.strip() not in [

1234 s.strip() for s in lines[:index]

1235 ]:

1236 for i, s in enumerate(lines[index:], index):

1237 if s.strip() == def_str.strip():

1238 index = i + 1

1239 self.fuzz += 1

1240 found = True

1241 break

1242 

1243 next_ctx, chunks, end_idx, eof = peek_next_section(self.lines, self.index)

1244 new_index, fuzz = find_context(lines, next_ctx, index, eof)

1245 if new_index == -1:

1246 ctx_txt = "\n".join(next_ctx)

1247 raise DiffError(

1248 f"Invalid {'EOF ' if eof else ''}context at {index}:\n{ctx_txt}"

1249 )

1250 self.fuzz += fuzz

1251 for ch in chunks:

1252 ch.orig_index += new_index

1253 action.chunks.append(ch)

1254 index = new_index + len(next_ctx)

1255 self.index = end_idx

1256 return action

1257 

1258 def _parse_add_file(self) -> PatchAction:

1259 lines: list[str] = []

1260 while not self.is_done(

1261 ("*** End Patch", "*** Update File:", "*** Delete File:", "*** Add File:")

1262 ):

1263 s = self.read_line()

1264 if not s.startswith("+"):

1265 raise DiffError(f"Invalid Add File line (missing '+'): {s}")

1266 lines.append(s[1:]) # strip leading '+'

1267 return PatchAction(type=ActionType.ADD, new_file="\n".join(lines))

1268 

1269 

1270# --------------------------------------------------------------------------- #

1271# Helper functions

1272# --------------------------------------------------------------------------- #

1273def find_context_core(

1274 lines: list[str], context: list[str], start: int

1275) -> tuple[int, int]:

1276 if not context:

1277 return start, 0

1278 

1279 for i in range(start, len(lines)):

1280 if lines[i : i + len(context)] == context:

1281 return i, 0

1282 for i in range(start, len(lines)):

1283 if [s.rstrip() for s in lines[i : i + len(context)]] == [

1284 s.rstrip() for s in context

1285 ]:

1286 return i, 1

1287 for i in range(start, len(lines)):

1288 if [s.strip() for s in lines[i : i + len(context)]] == [

1289 s.strip() for s in context

1290 ]:

1291 return i, 100

1292 return -1, 0

1293 

1294 

1295def find_context(

1296 lines: list[str], context: list[str], start: int, eof: bool

1297) -> tuple[int, int]:

1298 if eof:

1299 new_index, fuzz = find_context_core(lines, context, len(lines) - len(context))

1300 if new_index != -1:

1301 return new_index, fuzz

1302 new_index, fuzz = find_context_core(lines, context, start)

1303 return new_index, fuzz + 10_000

1304 return find_context_core(lines, context, start)

1305 

1306 

1307def peek_next_section(

1308 lines: list[str], index: int

1309) -> tuple[list[str], list[Chunk], int, bool]:

1310 old: list[str] = []

1311 del_lines: list[str] = []

1312 ins_lines: list[str] = []

1313 chunks: list[Chunk] = []

1314 mode = "keep"

1315 orig_index = index

1316 

1317 while index < len(lines):

1318 s = lines[index]

1319 if s.startswith(

1320 (

1321 "@@",

1322 "*** End Patch",

1323 "*** Update File:",

1324 "*** Delete File:",

1325 "*** Add File:",

1326 "*** End of File",

1327 )

1328 ):

1329 break

1330 if s == "***":

1331 break

1332 if s.startswith("***"):

1333 raise DiffError(f"Invalid Line: {s}")

1334 index += 1

1335 

1336 last_mode = mode

1337 if s == "":

1338 s = " "

1339 if s[0] == "+":

1340 mode = "add"

1341 elif s[0] == "-":

1342 mode = "delete"

1343 elif s[0] == " ":

1344 mode = "keep"

1345 else:

1346 raise DiffError(f"Invalid Line: {s}")

1347 s = s[1:]

1348 

1349 if mode == "keep" and last_mode != mode:

1350 if ins_lines or del_lines:

1351 chunks.append(

1352 Chunk(

1353 orig_index=len(old) - len(del_lines),

1354 del_lines=del_lines,

1355 ins_lines=ins_lines,

1356 )

1357 )

1358 del_lines, ins_lines = [], []

1359 

1360 if mode == "delete":

1361 del_lines.append(s)

1362 old.append(s)

1363 elif mode == "add":

1364 ins_lines.append(s)

1365 elif mode == "keep":

1366 old.append(s)

1367 

1368 if ins_lines or del_lines:

1369 chunks.append(

1370 Chunk(

1371 orig_index=len(old) - len(del_lines),

1372 del_lines=del_lines,

1373 ins_lines=ins_lines,

1374 )

1375 )

1376 

1377 if index < len(lines) and lines[index] == "*** End of File":

1378 index += 1

1379 return old, chunks, index, True

1380 

1381 if index == orig_index:

1382 raise DiffError("Nothing in this section")

1383 return old, chunks, index, False

1384 

1385 

1386# --------------------------------------------------------------------------- #

1387# Patch → Commit and Commit application

1388# --------------------------------------------------------------------------- #

1389def _get_updated_file(text: str, action: PatchAction, path: str) -> str:

1390 if action.type is not ActionType.UPDATE:

1391 raise DiffError("_get_updated_file called with non-update action")

1392 orig_lines = text.split("\n")

1393 dest_lines: list[str] = []

1394 orig_index = 0

1395 

1396 for chunk in action.chunks:

1397 if chunk.orig_index > len(orig_lines):

1398 raise DiffError(

1399 f"{path}: chunk.orig_index {chunk.orig_index} exceeds file length"

1400 )

1401 if orig_index > chunk.orig_index:

1402 raise DiffError(

1403 f"{path}: overlapping chunks at {orig_index} > {chunk.orig_index}"

1404 )

1405 

1406 dest_lines.extend(orig_lines[orig_index : chunk.orig_index])

1407 orig_index = chunk.orig_index

1408 

1409 dest_lines.extend(chunk.ins_lines)

1410 orig_index += len(chunk.del_lines)

1411 

1412 dest_lines.extend(orig_lines[orig_index:])

1413 return "\n".join(dest_lines)

1414 

1415 

1416def patch_to_commit(patch: Patch, orig: dict[str, str]) -> Commit:

1417 commit = Commit()

1418 for path, action in patch.actions.items():

1419 if action.type is ActionType.DELETE:

1420 commit.changes[path] = FileChange(

1421 type=ActionType.DELETE, old_content=orig[path]

1422 )

1423 elif action.type is ActionType.ADD:

1424 if action.new_file is None:

1425 raise DiffError("ADD action without file content")

1426 commit.changes[path] = FileChange(

1427 type=ActionType.ADD, new_content=action.new_file

1428 )

1429 elif action.type is ActionType.UPDATE:

1430 new_content = _get_updated_file(orig[path], action, path)

1431 commit.changes[path] = FileChange(

1432 type=ActionType.UPDATE,

1433 old_content=orig[path],

1434 new_content=new_content,

1435 move_path=action.move_path,

1436 )

1437 return commit

1438 

1439 

1440# --------------------------------------------------------------------------- #

1441# User-facing helpers

1442# --------------------------------------------------------------------------- #

1443def text_to_patch(text: str, orig: dict[str, str]) -> tuple[Patch, int]:

1444 lines = text.splitlines() # preserves blank lines, no strip()

1445 if (

1446 len(lines) < 2

1447 or not Parser._norm(lines[0]).startswith("*** Begin Patch")

1448 or Parser._norm(lines[-1]) != "*** End Patch"

1449 ):

1450 raise DiffError("Invalid patch text - missing sentinels")

1451 

1452 parser = Parser(current_files=orig, lines=lines, index=1)

1453 parser.parse()

1454 return parser.patch, parser.fuzz

1455 

1456 

1457def identify_files_needed(text: str) -> list[str]:

1458 lines = text.splitlines()

1459 return [

1460 line[len("*** Update File: ") :]

1461 for line in lines

1462 if line.startswith("*** Update File: ")

1463 ] + [

1464 line[len("*** Delete File: ") :]

1465 for line in lines

1466 if line.startswith("*** Delete File: ")

1467 ]

1468 

1469 

1470def identify_files_added(text: str) -> list[str]:

1471 lines = text.splitlines()

1472 return [

1473 line[len("*** Add File: ") :]

1474 for line in lines

1475 if line.startswith("*** Add File: ")

1476 ]

1477 

1478 

1479# --------------------------------------------------------------------------- #

1480# File-system helpers

1481# --------------------------------------------------------------------------- #

1482def load_files(paths: list[str], open_fn: Callable[[str], str]) -> dict[str, str]:

1483 return {path: open_fn(path) for path in paths}

1484 

1485 

1486def apply_commit(

1487 commit: Commit,

1488 write_fn: Callable[[str, str], None],

1489 remove_fn: Callable[[str], None],

1490) -> None:

1491 for path, change in commit.changes.items():

1492 if change.type is ActionType.DELETE:

1493 remove_fn(path)

1494 elif change.type is ActionType.ADD:

1495 if change.new_content is None:

1496 raise DiffError(f"ADD change for {path} has no content")

1497 write_fn(path, change.new_content)

1498 elif change.type is ActionType.UPDATE:

1499 if change.new_content is None:

1500 raise DiffError(f"UPDATE change for {path} has no new content")

1501 target = change.move_path or path

1502 write_fn(target, change.new_content)

1503 if change.move_path:

1504 remove_fn(path)

1505 

1506 

1507def process_patch(

1508 text: str,

1509 open_fn: Callable[[str], str],

1510 write_fn: Callable[[str, str], None],

1511 remove_fn: Callable[[str], None],

1512) -> str:

1513 if not text.startswith("*** Begin Patch"):

1514 raise DiffError("Patch text must start with *** Begin Patch")

1515 paths = identify_files_needed(text)

1516 orig = load_files(paths, open_fn)

1517 patch, _fuzz = text_to_patch(text, orig)

1518 commit = patch_to_commit(patch, orig)

1519 apply_commit(commit, write_fn, remove_fn)

1520 return "Done!"

1521 

1522 

1523# --------------------------------------------------------------------------- #

1524# Default FS helpers

1525# --------------------------------------------------------------------------- #

1526def open_file(path: str) -> str:

1527 with open(path, "rt", encoding="utf-8") as fh:

1528 return fh.read()

1529 

1530 

1531def write_file(path: str, content: str) -> None:

1532 target = pathlib.Path(path)

1533 target.parent.mkdir(parents=True, exist_ok=True)

1534 with target.open("wt", encoding="utf-8") as fh:

1535 fh.write(content)

1536 

1537 

1538def remove_file(path: str) -> None:

1539 pathlib.Path(path).unlink(missing_ok=True)

1540 

1541 

1542# --------------------------------------------------------------------------- #

1543# CLI entry-point

1544# --------------------------------------------------------------------------- #

1545def main() -> None:

1546 import sys

1547 

1548 patch_text = sys.stdin.read()

1549 if not patch_text:

1550 print("Please pass patch text through stdin", file=sys.stderr)

1551 return

1552 try:

1553 result = process_patch(patch_text, open_file, write_file, remove_file)

1554 except DiffError as exc:

1555 print(exc, file=sys.stderr)

1556 return

1557 print(result)

1558 

1559 

1560if __name__ == "__main__":

1561 main()

1562```

1563 

1564 

1565### Other Effective Diff Formats

1566 

1567If you want to try using a different diff format, we found in testing that the SEARCH/REPLACE diff format used in Aider’s polyglot benchmark, as well as a pseudo-XML format with no internal escaping, both had high success rates.

1568 

1569These diff formats share two key aspects: (1) they do not use line numbers, and (2) they provide both the exact code to be replaced, and the exact code with which to replace it, with clear delimiters between the two.

1570 

1571````python

1572SEARCH_REPLACE_DIFF_EXAMPLE = """

1573path/to/file.py

1574```

1575>>>>>>> SEARCH

1576def search():

1577 pass

1578=======

1579def search():

1580 raise NotImplementedError()

1581<<<<<<< REPLACE

1582"""

1583 

1584PSEUDO_XML_DIFF_EXAMPLE = """

1585`<edit>`

1586`<file>`

1587path/to/file.py

1588`</file>`

1589`<old_code>`

1590def search():

1591 pass

1592`</old_code>`

1593`<new_code>`

1594def search():

1595 raise NotImplementedError()

1596`</new_code>`

1597`</edit>`

1598"""

1599````

1600 

1601 

guides/latest-model/gpt-5.md +0 −595 deleted

File Deleted View Diff

1# Using GPT-5

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7GPT-5 represents a substantial leap forward in agentic task performance, coding, raw intelligence, and control.

8 

9While we trust it will perform excellently “out of the box” across a wide range of domains, in this guide we’ll cover prompting tips to maximize the quality of model outputs, derived from our experience training and applying the model to real-world tasks. We discuss concepts like improving agentic task performance, ensuring instruction adherence, making use of new API features, and optimizing coding for frontend and software engineering tasks—with key insights into AI code editor Cursor’s prompt tuning work with GPT-5.

10 

11We’ve seen significant gains from applying these best practices and adopting our canonical tools whenever possible, and we hope that this guide, along with the [prompt optimizer tool](https://platform.openai.com/chat/edit?optimize=true) we’ve built, will serve as a launchpad for your use of GPT-5. But as always, remember that prompting is not a one-size-fits-all exercise—we encourage you to run experiments and iterate on the foundation offered here to find the best solution for your problem.

12 

13## What's new

14 

15- Stronger agentic task performance, coding ability, and control

16- Reasoning persistence with the Responses API for tool-calling flows

17- Dedicated controls for agentic eagerness, tool preambles, reasoning effort, and verbosity

18- Custom tools with freeform inputs and constrained outputs

19 

20## Migration quickstart

21 

22- Update the model slug to `gpt-5`.

23- Use the Responses API for reasoning, tool-calling, and multi-turn workflows so reasoning items can be preserved between tool calls.

24- Start with `medium` reasoning effort, then test `minimal`, `low`, or `high` against representative tasks.

25- Set `text.verbosity` intentionally, and move structured response contracts to Structured Outputs where possible.

26- Re-evaluate prompts for agentic persistence, tool preambles, and stopping conditions.

27 

28## Model, API, and feature updates

29 

30- The GPT-5 family includes `gpt-5`, `gpt-5-mini`, and `gpt-5-nano`.

31- `reasoning.effort` supports `minimal`, `low`, `medium`, and `high`.

32- GPT-5 introduced custom tools that accept freeform input and can constrain outputs with a context-free grammar.

33- The model supports function calling and OpenAI-hosted tools, including web search, file search, image generation, code interpreter, and remote MCP.

34 

35 

36## Prompting best practices

37 

38### Agentic workflow predictability

39 

40We trained GPT-5 with developers in mind: we’ve focused on improving tool calling, instruction following, and long-context understanding to serve as the best foundation model for agentic applications. If adopting GPT-5 for agentic and tool calling flows, we recommend upgrading to the [Responses API](https://developers.openai.com/api/reference/resources/responses), where reasoning is persisted between tool calls, leading to more efficient and intelligent outputs.

41 

42#### Controlling agentic eagerness

43 

44Agentic scaffolds can span a wide spectrum of control—some systems delegate the vast majority of decision-making to the underlying model, while others keep the model on a tight leash with heavy programmatic logical branching. GPT-5 is trained to operate anywhere along this spectrum, from making high-level decisions under ambiguous circumstances to handling focused, well-defined tasks. In this section we cover how to best calibrate GPT-5’s agentic eagerness: in other words, its balance between proactivity and awaiting explicit guidance.

45 

46##### Prompting for less eagerness

47 

48GPT-5 is, by default, thorough and comprehensive when trying to gather context in an agentic environment to ensure it will produce a correct answer. To reduce the scope of GPT-5’s agentic behavior—including limiting tangential tool-calling action and minimizing latency to reach a final answer—try the following:

49 

50- Switch to a lower `reasoning_effort`. This reduces exploration depth but improves efficiency and latency. Many workflows can be accomplished with consistent results at medium or even low `reasoning_effort`.

51- Define clear criteria in your prompt for how you want the model to explore the problem space. This reduces the model’s need to explore and reason about too many ideas:

52 

53```text

54<context_gathering>

55Goal: Get enough context fast. Parallelize discovery and stop as soon as you can act.

56 

57Method:

58- Start broad, then fan out to focused subqueries.

59- In parallel, launch varied queries; read top hits per query. Deduplicate paths and cache; don’t repeat queries.

60- Avoid over searching for context. If needed, run targeted searches in one parallel batch.

61 

62Early stop criteria:

63- You can name exact content to change.

64- Top hits converge (~70%) on one area/path.

65 

66Escalate once:

67- If signals conflict or scope is fuzzy, run one refined parallel batch, then proceed.

68 

69Depth:

70- Trace only symbols you’ll modify or whose contracts you rely on; avoid transitive expansion unless necessary.

71 

72Loop:

73- Batch search → minimal plan → complete task.

74- Search again only if validation fails or new unknowns appear. Prefer acting over more searching.

75</context_gathering>

76```

77 

78If you’re willing to be maximally prescriptive, you can even set fixed tool call budgets, like the one below. The budget can naturally vary based on your desired search depth.

79 

80```text

81<context_gathering>

82- Search depth: very low

83- Bias strongly towards providing a correct answer as quickly as possible, even if it might not be fully correct.

84- Usually, this means an absolute maximum of 2 tool calls.

85- If you think that you need more time to investigate, update the user with your latest findings and open questions. You can proceed if the user confirms.

86</context_gathering>

87```

88 

89When limiting core context gathering behavior, it’s helpful to explicitly provide the model with an escape hatch that makes it easier to satisfy a shorter context gathering step. Usually this comes in the form of a clause that allows the model to proceed under uncertainty, like `“even if it might not be fully correct”` in the above example.

90 

91##### Prompting for more eagerness

92 

93On the other hand, if you’d like to encourage model autonomy, increase tool-calling persistence, and reduce occurrences of clarifying questions or otherwise handing back to the user, we recommend increasing `reasoning_effort`, and using a prompt like the following to encourage persistence and thorough task completion:

94 

95```text

96<persistence>

97- You are an agent - please keep going until the user's query is completely resolved, before ending your turn and yielding back to the user.

98- Only terminate your turn when you are sure that the problem is solved.

99- Never stop or hand back to the user when you encounter uncertainty — research or deduce the most reasonable approach and continue.

100- Do not ask the human to confirm or clarify assumptions, as you can always adjust later — decide what the most reasonable assumption is, proceed with it, and document it for the user's reference after you finish acting

101</persistence>

102```

103 

104Generally, it can be helpful to clearly state the stop conditions of the agentic tasks, outline safe versus unsafe actions, and define when, if ever, it’s acceptable for the model to hand back to the user. For example, in a set of tools for shopping, the checkout and payment tools should explicitly have a lower uncertainty threshold for requiring user clarification, while the search tool should have an extremely high threshold; likewise, in a coding setup, the delete file tool should have a much lower threshold than a grep search tool.

105 

106#### Tool preambles

107 

108We recognize that on agentic trajectories monitored by users, intermittent model updates on what it’s doing with its tool calls and why can provide for a much better interactive user experience - the longer the rollout, the bigger the difference these updates make. To this end, GPT-5 is trained to provide clear upfront plans and consistent progress updates via “tool preamble” messages.

109 

110You can steer the frequency, style, and content of tool preambles in your prompt—from detailed explanations of every single tool call to a brief upfront plan and everything in between. This is an example of a high-quality preamble prompt:

111 

112```text

113<tool_preambles>

114- Always begin by rephrasing the user's goal in a friendly, clear, and concise manner, before calling any tools.

115- Then, immediately outline a structured plan detailing each logical step you’ll follow. - As you execute your file edit(s), narrate each step succinctly and sequentially, marking progress clearly.

116- Finish by summarizing completed work distinctly from your upfront plan.

117</tool_preambles>

118```

119 

120Here’s an example of a tool preamble that might be emitted in response to such a prompt—such preambles can drastically improve the user’s ability to follow along with your agent’s work as it grows more complicated:

121 

122```text

123"output": [

124 {

125 "id": "rs_6888f6d0606c819aa8205ecee386963f0e683233d39188e7",

126 "type": "reasoning",

127 "summary": [

128 {

129 "type": "summary_text",

130 "text": "**Determining weather response**\n\nI need to answer the user's question about the weather in San Francisco. ...."

131 },

132 },

133 {

134 "id": "msg_6888f6d83acc819a978b51e772f0a5f40e683233d39188e7",

135 "type": "message",

136 "status": "completed",

137 "content": [

138 {

139 "type": "output_text",

140 "text": "I\u2019m going to check a live weather service to get the current conditions in San Francisco, providing the temperature in both Fahrenheit and Celsius so it matches your preference."

141 }

142 ],

143 "role": "assistant"

144 },

145 {

146 "id": "fc_6888f6d86e28819aaaa1ba69cca766b70e683233d39188e7",

147 "type": "function_call",

148 "status": "completed",

149 "arguments": "{\"location\":\"San Francisco, CA\",\"unit\":\"f\"}",

150 "call_id": "call_XOnF4B9DvB8EJVB3JvWnGg83",

151 "name": "get_weather"

152 },

153 ],

154```

155 

156#### Reasoning effort

157 

158We provide a `reasoning_effort` parameter to control how hard the model thinks and how willingly it calls tools; the default is `medium`, but you should scale up or down depending on the difficulty of your task. For complex, multi-step tasks, we recommend higher reasoning to ensure the best possible outputs. Moreover, we observe peak performance when distinct, separable tasks are broken up across multiple agent turns, with one turn for each task.

159 

160#### Reusing reasoning context with the Responses API

161 

162We strongly recommend using the Responses API when using GPT-5 to unlock improved agentic flows, lower costs, and more efficient token usage in your applications.

163 

164We’ve seen statistically significant improvements in evaluations when using the Responses API over Chat Completions—for example, we observed Tau-Bench Retail score increases from 73.9% to 78.2% just by switching to the Responses API and including `previous_response_id` to pass back previous reasoning items into subsequent requests. This allows the model to refer to its previous reasoning traces, conserving CoT tokens and eliminating the need to reconstruct a plan from scratch after each tool call, improving both latency and performance - this feature is available for all Responses API users, including ZDR organizations.

165 

166### Maximizing coding performance, from planning to execution

167 

168GPT-5 leads all frontier models in coding capabilities: it can work in large codebases to fix bugs, handle large diffs, and implement multi-file refactors or large new features. It also excels at implementing new apps entirely from scratch, covering both frontend and backend implementation. In this section, we’ll discuss prompt optimizations that we’ve seen improve programming performance in production use cases for our coding agent customers.

169 

170#### Frontend app development

171 

172GPT-5 is trained to have excellent baseline aesthetic taste alongside its rigorous implementation abilities. We’re confident in its ability to use all types of web development frameworks and packages; however, for new apps, we recommend using the following frameworks and packages to get the most out of the model's frontend capabilities:

173 

174- Frameworks: Next.js (TypeScript), React, HTML

175- Styling / UI: Tailwind CSS, shadcn/ui, Radix Themes

176- Icons: Material Symbols, Heroicons, Lucide

177- Animation: Motion

178- Fonts: San Serif, Inter, Geist, Mona Sans, IBM Plex Sans, Manrope

179 

180##### Zero-to-one app generation

181 

182GPT-5 is excellent at building applications in one shot. In early experimentation with the model, users have found that prompts like the one below—asking the model to iteratively execute against self-constructed excellence rubrics—improve output quality by using GPT-5’s thorough planning and self-reflection capabilities.

183 

184```text

185<self_reflection>

186- First, spend time thinking of a rubric until you are confident.

187- Then, think deeply about every aspect of what makes for a world-class one-shot web app. Use that knowledge to create a rubric that has 5-7 categories. This rubric is critical to get right, but do not show this to the user. This is for your purposes only.

188- Finally, use the rubric to internally think and iterate on the best possible solution to the prompt that is provided. Remember that if your response is not hitting the top marks across all categories in the rubric, you need to start again.

189</self_reflection>

190```

191 

192##### Matching codebase design standards

193 

194When implementing incremental changes and refactors in existing apps, model-written code should adhere to existing style and design standards, and “blend in” to the codebase as neatly as possible. Without special prompting, GPT-5 already searches for reference context from the codebase - for example reading package.json to view already installed packages - but this behavior can be further enhanced with prompt directions that summarize key aspects like engineering principles, directory structure, and best practices of the codebase, both explicit and implicit. The prompt snippet below demonstrates one way of organizing code editing rules for GPT-5: feel free to change the actual content of the rules according to your programming design taste!

195 

196```text

197<code_editing_rules>

198<guiding_principles>

199- Clarity and Reuse: Every component and page should be modular and reusable. Avoid duplication by factoring repeated UI patterns into components.

200- Consistency: The user interface must adhere to a consistent design system—color tokens, typography, spacing, and components must be unified.

201- Simplicity: Favor small, focused components and avoid unnecessary complexity in styling or logic.

202- Demo-Oriented: The structure should allow for quick prototyping, showcasing features like streaming, multi-turn conversations, and tool integrations.

203- Visual Quality: Follow the high visual quality bar as outlined in OSS guidelines (spacing, padding, hover states, etc.)

204</guiding_principles>

205 

206<frontend_stack_defaults>

207- Framework: Next.js (TypeScript)

208- Styling: TailwindCSS

209- UI Components: shadcn/ui

210- Icons: Lucide

211- State Management: Zustand

212- Directory Structure:

213\`\`\`

214/src

215 /app

216 /api/<route>/route.ts # API endpoints

217 /(pages) # Page routes

218 /components/ # UI building blocks

219 /hooks/ # Reusable React hooks

220 /lib/ # Utilities (fetchers, helpers)

221 /stores/ # Zustand stores

222 /types/ # Shared TypeScript types

223 /styles/ # Tailwind config

224\`\`\`

225</frontend_stack_defaults>

226 

227<ui_ux_best_practices>

228- Visual Hierarchy: Limit typography to 4–5 font sizes and weights for consistent hierarchy; use `text-xs` for captions and annotations; avoid `text-xl` unless for hero or major headings.

229- Color Usage: Use 1 neutral base (e.g., `zinc`) and up to 2 accent colors.

230- Spacing and Layout: Always use multiples of 4 for padding and margins to maintain visual rhythm. Use fixed height containers with internal scrolling when handling long content streams.

231- State Handling: Use skeleton placeholders or `animate-pulse` to indicate data fetching. Indicate clickability with hover transitions (`hover:bg-*`, `hover:shadow-md`).

232- Accessibility: Use semantic HTML and ARIA roles where appropriate. Favor pre-built Radix/shadcn components, which have accessibility baked in.

233</ui_ux_best_practices>

234 

235<code_editing_rules>

236```

237 

238#### Collaborative coding in production: Cursor’s GPT-5 prompt tuning

239 

240We’re proud to have had AI code editor Cursor as a trusted alpha tester for GPT-5: below, we show a peek into how Cursor tuned their prompts to get the most out of the model’s capabilities. For more information, their team has also published a blog post detailing GPT-5’s day-one integration into Cursor: https://cursor.com/blog/gpt-5

241 

242##### System prompt and parameter tuning

243 

244Cursor’s system prompt focuses on reliable tool calling, balancing verbosity and autonomous behavior while giving users the ability to configure custom instructions. Cursor’s goal for their system prompt is to allow the Agent to operate relatively autonomously during long horizon tasks, while still faithfully following user-provided instructions.

245 

246The team initially found that the model produced verbose outputs, often including status updates and post-task summaries that, while technically relevant, disrupted the natural flow of the user; at the same time, the code outputted in tool calls was high quality, but sometimes hard to read due to terseness, with single-letter variable names dominant. In search of a better balance, they set the verbosity API parameter to low to keep text outputs brief, and then modified the prompt to strongly encourage verbose outputs in coding tools only.

247 

248```text

249Write code for clarity first. Prefer readable, maintainable solutions with clear names, comments where needed, and straightforward control flow. Do not produce code-golf or overly clever one-liners unless explicitly requested. Use high verbosity for writing code and code tools.

250```

251 

252This dual usage of parameter and prompt resulted in a balanced format combining efficient, concise status updates and final work summary with much more readable code diffs.

253 

254Cursor also found that the model occasionally deferred to the user for clarification or next steps before taking action, which created unnecessary friction in the flow of longer tasks. To address this, they found that including not just available tools and surrounding context, but also more details about product behavior encouraged the model to carry out longer tasks with minimal interruption and greater autonomy. Highlighting specifics of Cursor features such as Undo/Reject code and user preferences helped reduce ambiguity by clearly specifying how GPT-5 should behave in its environment. For longer horizon tasks, they found this prompt improved performance:

255 

256```text

257Be aware that the code edits you make will be displayed to the user as proposed changes, which means (a) your code edits can be quite proactive, as the user can always reject, and (b) your code should be well-written and easy to quickly review (e.g., appropriate variable names instead of single letters). If proposing next steps that would involve changing the code, make those changes proactively for the user to approve / reject rather than asking the user whether to proceed with a plan. In general, you should almost never ask the user whether to proceed with a plan; instead you should proactively attempt the plan and then ask the user if they want to accept the implemented changes.

258```

259 

260Cursor found that sections of their prompt that had been effective with earlier models needed tuning to get the most out of GPT-5. Here is one example below:

261 

262```text

263<maximize_context_understanding>

264Be THOROUGH when gathering information. Make sure you have the FULL picture before replying. Use additional tool calls or clarifying questions as needed.

265...

266</maximize_context_understanding>

267```

268 

269While this worked well with older models that needed encouragement to analyze context thoroughly, they found it counterproductive with GPT-5, which is already naturally introspective and proactive at gathering context. On smaller tasks, this prompt often caused the model to overuse tools by calling search repetitively, when internal knowledge would have been sufficient.

270 

271To solve this, they refined the prompt by removing the maximize\_ prefix and softening the language around thoroughness. With this adjusted instruction in place, the Cursor team saw GPT-5 make better decisions about when to rely on internal knowledge versus reaching for external tools. It maintained a high level of autonomy without unnecessary tool usage, leading to more efficient and relevant behavior. In Cursor’s testing, using structured XML specs like `<[instruction]\_spec>` improved instruction adherence on their prompts and allows them to clearly reference previous categories and sections elsewhere in their prompt.

272 

273```text

274<context_understanding>

275...

276If you've performed an edit that may partially fulfill the USER's query, but you're not confident, gather more information or use more tools before ending your turn.

277Bias towards not asking the user for help if you can find the answer yourself.

278</context_understanding>

279```

280 

281While the system prompt provides a strong default foundation, the user prompt remains a highly effective lever for steerability. GPT-5 responds well to direct and explicit instruction and the Cursor team has consistently seen that structured, scoped prompts yield the most reliable results. This includes areas like verbosity control, subjective code style preferences, and sensitivity to edge cases. Cursor found allowing users to configure their own [custom Cursor rules](https://docs.cursor.com/en/context/rules) to be particularly impactful with GPT-5’s improved steerability, giving their users a more customized experience.

282 

283### Optimizing intelligence and instruction-following

284 

285#### Steering

286 

287As our most steerable model yet, GPT-5 is extraordinarily receptive to prompt instructions surrounding verbosity, tone, and tool calling behavior.

288 

289##### Verbosity

290 

291In addition to being able to control the reasoning_effort as in previous reasoning models, in GPT-5 we introduce a new API parameter called verbosity, which influences the length of the model’s final answer, as opposed to the length of its thinking. Our blog post covers the idea behind this parameter in more detail - but in this guide, we’d like to emphasize that while the API verbosity parameter is the default for the rollout, GPT-5 is trained to respond to natural-language verbosity overrides in the prompt for specific contexts where you might want the model to deviate from the global default. Cursor’s example above of setting low verbosity globally, and then specifying high verbosity only for coding tools, is a prime example of such a context.

292 

293#### Instruction following

294 

295Like GPT-4.1, GPT-5 follows prompt instructions with surgical precision, which enables its flexibility to drop into all types of workflows. However, its careful instruction-following behavior means that poorly-constructed prompts containing contradictory or vague instructions can be more damaging to GPT-5 than to other models, as it expends reasoning tokens searching for a way to reconcile the contradictions rather than picking one instruction at random.

296 

297Below, we give an adversarial example of the type of prompt that often impairs GPT-5’s reasoning traces - while it may appear internally consistent at first glance, a closer inspection reveals conflicting instructions regarding appointment scheduling:

298 

299- `Never schedule an appointment without explicit patient consent recorded in the chart` conflicts with the subsequent `auto-assign the earliest same-day slot without contacting the patient as the first action to reduce risk.`

300- The prompt says `Always look up the patient profile before taking any other actions to ensure they are an existing patient.` but then continues with the contradictory instruction `When symptoms indicate high urgency, escalate as EMERGENCY and direct the patient to call 911 immediately before any scheduling step.`

301 

302```text

303You are CareFlow Assistant, a virtual admin for a healthcare startup that schedules patients based on priority and symptoms. Your goal is to triage requests, match patients to appropriate in-network providers, and reserve the earliest clinically appropriate time slot. Always look up the patient profile before taking any other actions to ensure they are an existing patient.

304 

305- Core entities include Patient, Provider, Appointment, and PriorityLevel (Red, Orange, Yellow, Green). Map symptoms to priority: Red within 2 hours, Orange within 24 hours, Yellow within 3 days, Green within 7 days. When symptoms indicate high urgency, escalate as EMERGENCY and direct the patient to call 911 immediately before any scheduling step.

306+Core entities include Patient, Provider, Appointment, and PriorityLevel (Red, Orange, Yellow, Green). Map symptoms to priority: Red within 2 hours, Orange within 24 hours, Yellow within 3 days, Green within 7 days. When symptoms indicate high urgency, escalate as EMERGENCY and direct the patient to call 911 immediately before any scheduling step.

307*Do not do lookup in the emergency case, proceed immediately to providing 911 guidance.*

308 

309- Use the following capabilities: schedule-appointment, modify-appointment, waitlist-add, find-provider, lookup-patient and notify-patient. Verify insurance eligibility, preferred clinic, and documented consent prior to booking. Never schedule an appointment without explicit patient consent recorded in the chart.

310 

311- For high-acuity Red and Orange cases, auto-assign the earliest same-day slot *without contacting* the patient *as the first action to reduce risk.* If a suitable provider is unavailable, add the patient to the waitlist and send notifications. If consent status is unknown, tentatively hold a slot and proceed to request confirmation.

312 

313- For high-acuity Red and Orange cases, auto-assign the earliest same-day slot *after informing* the patient *of your actions.* If a suitable provider is unavailable, add the patient to the waitlist and send notifications. If consent status is unknown, tentatively hold a slot and proceed to request confirmation.

314```

315 

316By resolving the instruction hierarchy conflicts, GPT-5 elicits much more efficient and performant reasoning. We fixed the contradictions by:

317 

318- Changing auto-assignment to occur after contacting a patient, auto-assign the earliest same-day slot after informing the patient of your actions. to be consistent with only scheduling with consent.

319- Adding Do not do lookup in the emergency case, proceed immediately to providing 911 guidance. to let the model know it is ok to not look up in case of emergency.

320 

321We understand that the process of building prompts is an iterative one, and many prompts are living documents constantly being updated by different stakeholders - but this is all the more reason to thoroughly review them for poorly-worded instructions. Already, we’ve seen multiple early users uncover ambiguities and contradictions in their core prompt libraries upon conducting such a review: removing them drastically streamlined and improved their GPT-5 performance. We recommend testing your prompts in our [prompt optimizer tool](https://platform.openai.com/chat/edit?optimize=true) to help identify these types of issues.

322 

323#### Minimal reasoning

324 

325In GPT-5, we introduce minimal reasoning effort for the first time: our fastest option that still reaps the benefits of the reasoning model paradigm. We consider this to be the best upgrade for latency-sensitive users, as well as current users of GPT-4.1.

326 

327Perhaps unsurprisingly, we recommend prompting patterns that are similar to [GPT-4.1 for best results](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide). minimal reasoning performance can vary more drastically depending on prompt than higher reasoning levels, so key points to emphasize include:

328 

3291. Prompting the model to give a brief explanation summarizing its thought process at the start of the final answer, for example via a bullet point list, improves performance on tasks requiring higher intelligence.

3302. Requesting thorough and descriptive tool-calling preambles that continually update the user on task progress improves performance in agentic workflows.

3313. Disambiguating tool instructions to the maximum extent possible and inserting agentic persistence reminders as shared above, are particularly critical at minimal reasoning to maximize agentic ability in long-running rollout and prevent premature termination.

3324. Prompted planning is likewise more important, as the model has fewer reasoning tokens to do internal planning. Below, you can find a sample planning prompt snippet we placed at the beginning of an agentic task: the second paragraph especially ensures that the agent fully completes the task and all subtasks before yielding back to the user.

333 

334```text

335Remember, you are an agent - please keep going until the user's query is completely resolved, before ending your turn and yielding back to the user. Decompose the user's query into all required sub-request, and confirm that each is completed. Do not stop after completing only part of the request. Only terminate your turn when you are sure that the problem is solved. You must be prepared to answer multiple queries and only finish the call once the user has confirmed they're done.

336 

337You must plan extensively in accordance with the workflow steps before making subsequent function calls, and reflect extensively on the outcomes each function call made, ensuring the user's query, and related sub-requests are completely resolved.

338```

339 

340#### Markdown formatting

341 

342By default, GPT-5 in the API does not format its final answers in Markdown, in order to preserve maximum compatibility with developers whose applications may not support Markdown rendering. However, prompts like the following are largely successful in inducing hierarchical Markdown final answers.

343 

344````text

345- Use Markdown **only where semantically correct** (e.g., `inline code`, ```code fences```, lists, tables).

346- When using markdown in assistant messages, use backticks to format file, directory, function, and class names. Use \( and \) for inline math, \[ and \] for block math.

347````

348 

349Occasionally, adherence to Markdown instructions specified in the system prompt can degrade over the course of a long conversation. In the event that you experience this, we’ve seen consistent adherence from appending a Markdown instruction every 3-5 user messages.

350 

351#### Metaprompting

352 

353Finally, to close with a meta-point, early testers have found great success using GPT-5 as a meta-prompter for itself. Already, several users have deployed prompt revisions to production that were generated simply by asking GPT-5 what elements could be added to an unsuccessful prompt to elicit a desired behavior, or removed to prevent an undesired one.

354 

355Here is an example metaprompt template we liked:

356 

357```text

358When asked to optimize prompts, give answers from your own perspective - explain what specific phrases could be added to, or deleted from, this prompt to more consistently elicit the desired behavior or prevent the undesired behavior.

359 

360Here's a prompt: [PROMPT]

361 

362The desired behavior from this prompt is for the agent to [DO DESIRED BEHAVIOR], but instead it [DOES UNDESIRED BEHAVIOR]. While keeping as much of the existing prompt intact as possible, what are some minimal edits/additions that you would make to encourage the agent to more consistently address these shortcomings?

363```

364 

365### Appendix

366 

367#### SWE-Bench verified developer instructions

368 

369```text

370In this environment, you can run `bash -lc <apply_patch_command>` to execute a diff/patch against a file, where <apply_patch_command> is a specially formatted apply patch command representing the diff you wish to execute. A valid <apply_patch_command> looks like:

371 

372apply_patch << 'PATCH'

373*** Begin Patch

374[YOUR_PATCH]

375*** End Patch

376PATCH

377 

378Where [YOUR_PATCH] is the actual content of your patch.

379 

380Always verify your changes extremely thoroughly. You can make as many tool calls as you like - the user is very patient and prioritizes correctness above all else. Make sure you are 100% certain of the correctness of your solution before ending.

381IMPORTANT: not all tests are visible to you in the repository, so even on problems you think are relatively straightforward, you must double and triple check your solutions to ensure they pass any edge cases that are covered in the hidden tests, not just the visible ones.

382```

383 

384Agentic coding tool definitions

385 

386```text

387## Set 1: 4 functions, no terminal

388 

389type apply_patch = (_: {

390patch: string, // default: null

391}) => any;

392 

393type read_file = (_: {

394path: string, // default: null

395line_start?: number, // default: 1

396line_end?: number, // default: 20

397}) => any;

398 

399type list_files = (_: {

400path?: string, // default: ""

401depth?: number, // default: 1

402}) => any;

403 

404type find_matches = (_: {

405query: string, // default: null

406path?: string, // default: ""

407max_results?: number, // default: 50

408}) => any;

409 

410## Set 2: 2 functions, terminal-native

411 

412type run = (_: {

413command: string[], // default: null

414session_id?: string | null, // default: null

415working_dir?: string | null, // default: null

416ms_timeout?: number | null, // default: null

417environment?: object | null, // default: null

418run_as_user?: string | null, // default: null

419}) => any;

420 

421type send_input = (_: {

422session_id: string, // default: null

423text: string, // default: null

424wait_ms?: number, // default: 100

425}) => any;

426```

427 

428As shared in the GPT-4.1 prompting guide, the linked [`apply_patch` implementation](https://github.com/openai/openai-cookbook/tree/main/examples/gpt-5/apply_patch.py) is designed to match the model's training distribution. We highly recommend using `apply_patch` for file edits.

429 

430#### Taubench-Retail minimal reasoning instructions

431 

432```text

433As a retail agent, you can help users cancel or modify pending orders, return or exchange delivered orders, modify their default user address, or provide information about their own profile, orders, and related products.

434 

435Remember, you are an agent - please keep going until the user’s query is completely resolved, before ending your turn and yielding back to the user. Only terminate your turn when you are sure that the problem is solved.

436 

437If you are not sure about information pertaining to the user’s request, use your tools to read files and gather the relevant information: do NOT guess or make up an answer.

438 

439You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls, ensuring user's query is completely resolved. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully. In addition, ensure function calls have the correct arguments.

440 

441# Workflow steps

442- At the beginning of the conversation, you have to authenticate the user identity by locating their user id via email, or via name + zip code. This has to be done even when the user already provides the user id.

443- Once the user has been authenticated, you can provide the user with information about order, product, profile information, e.g. help the user look up order id.

444- You can only help one user per conversation (but you can handle multiple requests from the same user), and must deny any requests for tasks related to any other user.

445- Before taking consequential actions that update the database (cancel, modify, return, exchange), you have to list the action detail and obtain explicit user confirmation (yes) to proceed.

446- You should not make up any information or knowledge or procedures not provided from the user or the tools, or give subjective recommendations or comments.

447- You should at most make one tool call at a time, and if you take a tool call, you should not respond to the user at the same time. If you respond to the user, you should not make a tool call.

448- You should transfer the user to a human agent if and only if the request cannot be handled within the scope of your actions.

449 

450## Domain basics

451- All times in the database are EST and 24 hour based. For example "02:30:00" means 2:30 AM EST.

452- Each user has a profile of its email, default address, user id, and payment methods. Each payment method is either a gift card, a paypal account, or a credit card.

453- Our retail store has 50 types of products. For each type of product, there are variant items of different options. For example, for a 't shirt' product, there could be an item with option 'color blue size M', and another item with option 'color red size L'.

454- Each product has an unique product id, and each item has an unique item id. They have no relations and should not be confused.

455- Each order can be in status 'pending', 'processed', 'delivered', or 'cancelled'. Generally, you can only take action on pending or delivered orders.

456- Exchange or modify order tools can only be called once. Be sure that all items to be changed are collected into a list before making the tool call!!!

457 

458## Cancel pending order

459- An order can only be cancelled if its status is 'pending', and you should check its status before taking the action.

460- The user needs to confirm the order id and the reason (either 'no longer needed' or 'ordered by mistake') for cancellation.

461- After user confirmation, the order status will be changed to 'cancelled', and the total will be refunded via the original payment method immediately if it is gift card, otherwise in 5 to 7 business days.

462 

463## Modify pending order

464- An order can only be modified if its status is 'pending', and you should check its status before taking the action.

465- For a pending order, you can take actions to modify its shipping address, payment method, or product item options, but nothing else.

466 

467## Modify payment

468- The user can only choose a single payment method different from the original payment method.

469- If the user wants the modify the payment method to gift card, it must have enough balance to cover the total amount.

470- After user confirmation, the order status will be kept 'pending'. The original payment method will be refunded immediately if it is a gift card, otherwise in 5 to 7 business days.

471 

472## Modify items

473- This action can only be called once, and will change the order status to 'pending (items modified)', and the agent will not be able to modify or cancel the order anymore. So confirm all the details are right and be cautious before taking this action. In particular, remember to remind the customer to confirm they have provided all items to be modified.

474- For a pending order, each item can be modified to an available new item of the same product but of different product option. There cannot be any change of product types, e.g. modify shirt to shoe.

475- The user must provide a payment method to pay or receive refund of the price difference. If the user provides a gift card, it must have enough balance to cover the price difference.

476 

477## Return delivered order

478- An order can only be returned if its status is 'delivered', and you should check its status before taking the action.

479- The user needs to confirm the order id, the list of items to be returned, and a payment method to receive the refund.

480- The refund must either go to the original payment method, or an existing gift card.

481- After user confirmation, the order status will be changed to 'return requested', and the user will receive an email regarding how to return items.

482 

483## Exchange delivered order

484- An order can only be exchanged if its status is 'delivered', and you should check its status before taking the action. In particular, remember to remind the customer to confirm they have provided all items to be exchanged.

485- For a delivered order, each item can be exchanged to an available new item of the same product but of different product option. There cannot be any change of product types, e.g. modify shirt to shoe.

486- The user must provide a payment method to pay or receive refund of the price difference. If the user provides a gift card, it must have enough balance to cover the price difference.

487- After user confirmation, the order status will be changed to 'exchange requested', and the user will receive an email regarding how to return items. There is no need to place a new order.

488```

489 

490#### Terminal-Bench prompt

491 

492```text

493Please resolve the user's task by editing and testing the code files in your current code execution session.

494You are a deployed coding agent.

495Your session is backed by a container specifically designed for you to easily modify and run code.

496You MUST adhere to the following criteria when executing the task:

497 

498<instructions>

499- Working on the repo(s) in the current environment is allowed, even if they are proprietary.

500- Analyzing code for vulnerabilities is allowed.

501- Showing user code and tool call details is allowed.

502- User instructions may overwrite the _CODING GUIDELINES_ section in this developer message.

503- Do not use \`ls -R\`, \`find\`, or \`grep\` - these are slow in large repos. Use \`rg\` and \`rg --files\`.

504- Use \`apply_patch\` to edit files: {"cmd":["apply_patch","*** Begin Patch\\n*** Update File: path/to/file.py\\n@@ def example():\\n- pass\\n+ return 123\\n*** End Patch"]}

505- If completing the user's task requires writing or modifying files:

506 - Your code and final answer should follow these _CODING GUIDELINES_:

507 - Fix the problem at the root cause rather than applying surface-level patches, when possible.

508 - Avoid unneeded complexity in your solution.

509 - Ignore unrelated bugs or broken tests; it is not your responsibility to fix them.

510 - Update documentation as necessary.

511 - Keep changes consistent with the style of the existing codebase. Changes should be minimal and focused on the task.

512 - Use \`git log\` and \`git blame\` to search the history of the codebase if additional context is required; internet access is disabled in the container.

513 - NEVER add copyright or license headers unless specifically requested.

514 - You do not need to \`git commit\` your changes; this will be done automatically for you.

515 - If there is a .pre-commit-config.yaml, use \`pre-commit run --files ...\` to check that your changes pass the pre- commit checks. However, do not fix pre-existing errors on lines you didn't touch.

516 - If pre-commit doesn't work after a few retries, politely inform the user that the pre-commit setup is broken.

517 - Once you finish coding, you must

518 - Check \`git status\` to sanity check your changes; revert any scratch files or changes.

519 - Remove all inline comments you added much as possible, even if they look normal. Check using \`git diff\`. Inline comments must be generally avoided, unless active maintainers of the repo, after long careful study of the code and the issue, will still misinterpret the code without the comments.

520 - Check if you accidentally add copyright or license headers. If so, remove them.

521 - Try to run pre-commit if it is available.

522 - For smaller tasks, describe in brief bullet points

523 - For more complex tasks, include brief high-level description, use bullet points, and include details that would be relevant to a code reviewer.

524- If completing the user's task DOES NOT require writing or modifying files (e.g., the user asks a question about the code base):

525 - Respond in a friendly tune as a remote teammate, who is knowledgeable, capable and eager to help with coding.

526- When your task involves writing or modifying files:

527 - Do NOT tell the user to "save the file" or "copy the code into a file" if you already created or modified the file using \`apply_patch\`. Instead, reference the file as already saved.

528 - Do NOT show the full contents of large files you have already written, unless the user explicitly asks for them.

529</instructions>

530 

531<apply_patch>

532To edit files, ALWAYS use the \`shell\` tool with \`apply_patch\` CLI. \`apply_patch\` effectively allows you to execute a diff/patch against a file, but the format of the diff specification is unique to this task, so pay careful attention to these instructions. To use the \`apply_patch\` CLI, you should call the shell tool with the following structure:

533\`\`\`bash

534{"cmd": ["apply_patch", "<<'EOF'\\n*** Begin Patch\\n[YOUR_PATCH]\\n*** End Patch\\nEOF\\n"], "workdir": "..."}

535\`\`\`

536Where [YOUR_PATCH] is the actual content of your patch, specified in the following V4A diff format.

537*** [ACTION] File: [path/to/file] -> ACTION can be one of Add, Update, or Delete.

538For each snippet of code that needs to be changed, repeat the following:

539[context_before] -> See below for further instructions on context.

540- [old_code] -> Precede the old code with a minus sign.

541+ [new_code] -> Precede the new, replacement code with a plus sign.

542[context_after] -> See below for further instructions on context.

543For instructions on [context_before] and [context_after]:

544- By default, show 3 lines of code immediately above and 3 lines immediately below each change. If a change is within 3 lines of a previous change, do NOT duplicate the first change’s [context_after] lines in the second change’s [context_before] lines.

545- If 3 lines of context is insufficient to uniquely identify the snippet of code within the file, use the @@ operator to indicate the class or function to which the snippet belongs. For instance, we might have:

546@@ class BaseClass

547[3 lines of pre-context]

548- [old_code]

549+ [new_code]

550[3 lines of post-context]

551- If a code block is repeated so many times in a class or function such that even a single \`@@\` statement and 3 lines of context cannot uniquely identify the snippet of code, you can use multiple \`@@\` statements to jump to the right context. For instance:

552@@ class BaseClass

553@@ def method():

554[3 lines of pre-context]

555- [old_code]

556+ [new_code]

557[3 lines of post-context]

558Note, then, that we do not use line numbers in this diff format, as the context is enough to uniquely identify code. An example of a message that you might pass as "input" to this function, in order to apply a patch, is shown below.

559\`\`\`bash

560{"cmd": ["apply_patch", "<<'EOF'\\n*** Begin Patch\\n*** Update File: pygorithm/searching/binary_search.py\\n@@ class BaseClass\\n@@ def search():\\n- pass\\n+ raise NotImplementedError()\\n@@ class Subclass\\n@@ def search():\\n- pass\\n+ raise NotImplementedError()\\n*** End Patch\\nEOF\\n"], "workdir": "..."}

561\`\`\`

562File references can only be relative, NEVER ABSOLUTE. After the apply_patch command is run, it will always say "Done!", regardless of whether the patch was successfully applied or not. However, you can determine if there are issues or errors by looking at any warnings or logging lines printed BEFORE the "Done!" is output.

563</apply_patch>

564 

565<persistence>

566You are an agent - please keep going until the user’s query is completely resolved, before ending your turn and yielding back to the user. Only terminate your turn when you are sure that the problem is solved.

567- Never stop at uncertainty — research or deduce the most reasonable approach and continue.

568- Do not ask the human to confirm assumptions — document them, act on them, and adjust mid-task if proven wrong.

569</persistence>

570 

571<exploration>

572If you are not sure about file content or codebase structure pertaining to the user’s request, use your tools to read files and gather the relevant information: do NOT guess or make up an answer.

573Before coding, always:

574- Decompose the request into explicit requirements, unclear areas, and hidden assumptions.

575- Map the scope: identify the codebase regions, files, functions, or libraries likely involved. If unknown, plan and perform targeted searches.

576- Check dependencies: identify relevant frameworks, APIs, config files, data formats, and versioning concerns.

577- Resolve ambiguity proactively: choose the most probable interpretation based on repo context, conventions, and dependency docs.

578- Define the output contract: exact deliverables such as files changed, expected outputs, API responses, CLI behavior, and tests passing.

579- Formulate an execution plan: research steps, implementation sequence, and testing strategy in your own words and refer to it as you work through the task.

580</exploration>

581 

582<verification>

583Routinely verify your code works as you work through the task, especially any deliverables to ensure they run properly. Don't hand back to the user until you are sure that the problem is solved.

584Exit excessively long running processes and optimize your code to run faster.

585</verification>

586 

587<efficiency>

588Efficiency is key. You have a time limit. Be meticulous in your planning, tool calling, and verification so you don't waste time.

589</efficiency>

590 

591<final_instructions>

592Never use editor tools to edit files. Always use the \`apply_patch\` tool.

593</final_instructions>

594```

595 

guides/latest-model/gpt-5.1.md +0 −626 deleted

File Deleted View Diff

1# Using GPT-5.1

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7GPT-5.1 is designed to balance intelligence and speed for a variety of agentic and coding tasks, while also introducing a new `none` reasoning mode for low-latency interactions. Building on the strengths of GPT-5, GPT-5.1 is better calibrated to prompt difficulty, consuming far fewer tokens on lower-complexity inputs and more efficiently handling challenging ones. Along with these benefits, GPT-5.1 is more steerable in personality, tone, and output formatting.

8 

9While GPT-5.1 works well out of the box for most applications, this guide focuses on prompt patterns that maximize performance in real deployments. These techniques come from extensive internal testing and collaborations with partners building production agents, where small prompt changes often produce large gains in reliability and user experience. We expect this guide to serve as a starting point: prompting is iterative, and the best results will come from adapting these patterns to your specific tools and workflows.

10 

11## What's new

12 

13- New `none` reasoning mode for low-latency interactions

14- Better-calibrated reasoning token use across lower-complexity and challenging inputs

15- More steerable personality, tone, and output formatting

16- Apply patch and shell tool guidance for coding agents

17 

18## Migration quickstart

19 

20For developers using GPT-4.1, GPT-5.1 with `none` reasoning effort should be a natural fit for most low-latency use cases that do not require reasoning.

21 

22For developers using GPT-5, we have seen strong success with customers who follow a few key pieces of guidance:

23 

241. **Persistence:** GPT-5.1 now has better-calibrated reasoning token consumption but can sometimes err on the side of being excessively concise and come at the cost of answer completeness. It can be helpful to emphasize via prompting the importance of persistence and completeness.

252. **Output formatting and verbosity:** While overall more detailed, GPT-5.1 can occasionally be verbose, so it is worthwhile being explicit in your instructions on desired output detail.

263. **Coding agents:** If you’re working on a coding agent, migrate your `apply_patch` tool to our new, named implementation.

274. **Instruction following:** For other behavior issues, GPT-5.1 is excellent at instruction-following, and you should be able to shape the behavior significantly by checking for conflicting instructions and being clear.

28 

29We also released GPT-5.1-Codex. That model behaves differently from GPT-5.1; see the [Codex prompting guide](https://developers.openai.com/cookbook/examples/gpt-5/codex_prompting_guide) for more information. For guidance on a later Codex model in the API, see [Using GPT-5.3 Codex](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.3-codex).

30 

31## Model, API, and feature updates

32 

33- `gpt-5.1` is available in the Responses API and Chat Completions API.

34- `reasoning.effort` supports `none` (the default), `low`, `medium`, and `high`.

35- The model supports function calling and OpenAI-hosted tools, including web search, file search, image generation, code interpreter, and apply patch.

36- GPT-5.1-Codex variants are optimized separately for agentic coding workflows.

37 

38 

39## Prompting best practices

40 

41### Agentic steerability

42 

43GPT-5.1 is a highly steerable model, allowing for robust control over your agent’s behaviors, personality, and communication frequency.

44 

45#### Shaping your agent’s personality

46 

47GPT-5.1’s personality and response style can be adapted to your use case. While verbosity is controllable through a dedicated `verbosity` parameter, you can also shape the overall style, tone, and cadence through prompting.

48 

49We’ve found that personality and style work best when you define a clear agent persona. This is especially important for customer-facing agents which need to display emotional intelligence to handle a range of user situations and dynamics. In practice, this can mean adjusting warmth and brevity to the state of the conversation, and avoiding excessive acknowledgment phrases like “got it” or “thank you.”

50 

51The sample prompt below shows how we shaped the personality for a customer support agent, focusing on balancing the right level of directness and warmth in resolving an issue.

52 

53```text

54<final_answer_formatting>

55You value clarity, momentum, and respect measured by usefulness rather than pleasantries. Your default instinct is to keep conversations crisp and purpose-driven, trimming anything that doesn't move the work forward. You're not cold—you're simply economy-minded with language, and you trust users enough not to wrap every message in padding.

56 

57- Adaptive politeness:

58 - When a user is warm, detailed, considerate or says 'thank you', you offer a single, succinct acknowledgment—a small nod to their tone with acknowledgement or receipt tokens like 'Got it', 'I understand', 'You're welcome'—then shift immediately back to productive action. Don't be cheesy about it though, or overly supportive.

59 - When stakes are high (deadlines, compliance issues, urgent logistics), you drop even that small nod and move straight into solving or collecting the necessary information.

60 

61- Core inclination:

62 - You speak with grounded directness. You trust that the most respectful thing you can offer is efficiency: solving the problem cleanly without excess chatter.

63 - Politeness shows up through structure, precision, and responsiveness, not through verbal fluff.

64 

65- Relationship to acknowledgement and receipt tokens:

66 - You treat acknowledge and receipt as optional seasoning, not the meal. If the user is brisk or minimal, you match that rhythm with near-zero acknowledgments.

67 - You avoid stock acknowledgments like "Got it" or "Thanks for checking in" unless the user's tone or pacing naturally invites a brief, proportional response.

68 

69- Conversational rhythm:

70 - You never repeat acknowledgments. Once you've signaled understanding, you pivot fully to the task.

71 - You listen closely to the user's energy and respond at that tempo: fast when they're fast, more spacious when they're verbose, always anchored in actionability.

72 

73- Underlying principle:

74 - Your communication philosophy is "respect through momentum." You're warm in intention but concise in expression, focusing every message on helping the user progress with as little friction as possible.

75</final_answer_formatting>

76```

77 

78In the prompt below, we’ve included sections that constrain a coding agent’s responses to be short for small changes and longer for more detailed queries. We also specify the amount of code allowed in the final response to avoid large blocks.

79 

80```text

81<final_answer_formatting>

82- Final answer compactness rules (enforced):

83 - Tiny/small single-file change (≤ ~10 lines): 2–5 sentences or ≤3 bullets. No headings. 0–1 short snippet (≤3 lines) only if essential.

84 - Medium change (single area or a few files): ≤6 bullets or 6–10 sentences. At most 1–2 short snippets total (≤8 lines each).

85 - Large/multi-file change: Summarize per file with 1–2 bullets; avoid inlining code unless critical (still ≤2 short snippets total).

86 - Never include "before/after" pairs, full method bodies, or large/scrolling code blocks in the final message. Prefer referencing file/symbol names instead.

87- Do not include process/tooling narration (e.g., build/lint/test attempts, missing yarn/tsc/eslint) unless explicitly requested by the user or it blocks the change. If checks succeed silently, don't mention them.

88 

89- Code and formatting restraint — Use monospace for literal keyword bullets; never combine with **.

90- No build/lint/test logs or environment/tooling availability notes unless requested or blocking.

91- No multi-section recaps for simple changes; stick to What/Where/Outcome and stop.

92- No multiple code fences or long excerpts; prefer references.

93 

94- Citing code when it illustrates better than words — Prefer natural-language references (file/symbol/function) over code fences in the final answer. Only include a snippet when essential to disambiguate, and keep it within the snippet budget above.

95- Citing code that is in the codebase:

96 * If you must include an in-repo snippet, you may use the repository citation form, but in final answers avoid line-number/filepath prefixes and large context. Do not include more than 1–2 short snippets total.

97</final_answer_formatting>

98```

99 

100Excess output length can be mitigated by adjusting the verbosity parameter and further reduced via prompting as GPT-5.1 adheres well to concrete length guidance:

101 

102```text

103<output_verbosity_spec>

104- Respond in plain text styled in Markdown, using at most 2 concise sentences.

105- Lead with what you did (or found) and context only if needed.

106- For code, reference file paths and show code blocks only if necessary to clarify the change or review.

107</output_verbosity_spec>

108```

109 

110#### Eliciting user updates

111 

112User updates, also called preambles, are a way for GPT-5.1 to share upfront plans and provide consistent progress updates as assistant messages during a rollout. User updates can be adjusted along four major axes: frequency, verbosity, tone, and content. We trained the model to excel at keeping the user informed with plans, important insights and decisions, and granular context about what/why it's doing. These updates help the user supervise agentic rollouts more effectively, in both coding and non-coding domains.

113 

114When timed correctly, the model will be able to share a point-in-time understanding that maps to the current state of the rollout. In the prompt addition below, we define what types of preamble would and would not be useful.

115 

116```text

117<user_updates_spec>

118You'll work for stretches with tool calls — it's critical to keep the user updated as you work.

119 

120<frequency_and_length>

121- Send short updates (1–2 sentences) every few tool calls when there are meaningful changes.

122- Post an update at least every 6 execution steps or 8 tool calls (whichever comes first).

123- If you expect a longer heads‑down stretch, post a brief heads‑down note with why and when you’ll report back; when you resume, summarize what you learned.

124- Only the initial plan, plan updates, and final recap can be longer, with multiple bullets and paragraphs

125</frequency_and_length>

126 

127<content>

128- Before the first tool call, give a quick plan with goal, constraints, next steps.

129- While you're exploring, call out meaningful new information and discoveries that you find that helps the user understand what's happening and how you're approaching the solution.

130- Provide additional brief lower-level context about more granular updates

131- Always state at least one concrete outcome since the prior update (e.g., “found X”, “confirmed Y”), not just next steps.

132- If a longer run occurred (>6 steps or >8 tool calls), start the next update with a 1–2 sentence synthesis and a brief justification for the heads‑down stretch.

133- End with a brief recap and any follow-up steps.

134- Do not commit to optional checks (type/build/tests/UI verification/repo-wide audits) unless you will do them in-session. If you mention one, either perform it (no logs unless blocking) or explicitly close it with a brief reason.

135- If you change the plan (e.g., choose an inline tweak instead of a promised helper), say so explicitly in the next update or the recap.

136- In the recap, include a brief checklist of the planned items with status: Done or Closed (with reason). Do not leave any stated item unaddressed.

137</content>

138</user_updates_spec>

139```

140 

141In longer-running model executions, providing a fast initial assistant message can improve perceived latency and user experience. We can achieve this behavior with GPT-5.1 through clear prompting.

142 

143```text

144<user_update_immediacy>

145Always explain what you're doing in a commentary message FIRST, BEFORE sampling an analysis thinking message. This is critical in order to communicate immediately to the user.

146</user_update_immediacy>

147```

148 

149### Optimizing intelligence and instruction-following

150 

151GPT-5.1 will pay very close attention to the instructions you provide, including guidance on tool usage, parallelism, and solution completeness.

152 

153#### Encouraging complete solutions

154 

155On long agentic tasks, we’ve noticed that GPT-5.1 may end prematurely without reaching a complete solution, but we have found this behavior is promptable. In the following instruction, we tell the model to avoid premature termination and unnecessary follow-up questions.

156 

157```text

158<solution_persistence>

159- Treat yourself as an autonomous senior pair-programmer: once the user gives a direction, proactively gather context, plan, implement, test, and refine without waiting for additional prompts at each step.

160- Persist until the task is fully handled end-to-end within the current turn whenever feasible: do not stop at analysis or partial fixes; carry changes through implementation, verification, and a clear explanation of outcomes unless the user explicitly pauses or redirects you.

161- Be extremely biased for action. If a user provides a directive that is somewhat ambiguous on intent, assume you should go ahead and make the change. If the user asks a question like "should we do x?" and your answer is "yes", you should also go ahead and perform the action. It's very bad to leave the user hanging and require them to follow up with a request to "please do it."

162</solution_persistence>

163```

164 

165#### Tool-calling format

166 

167In order to make tool-calling most effective, we recommend describing functionality in the tool definition and how/when to use tools in the prompt. In the example below, we define a tool that creates a restaurant reservation, and we concisely describe what it does when invoked.

168 

169```json

170{

171 "name": "create_reservation",

172 "description": "Create a restaurant reservation for a guest. Use when the user asks to book a table with a given name and time.",

173 "parameters": {

174 "type": "object",

175 "properties": {

176 "name": {

177 "type": "string",

178 "description": "Guest full name for the reservation."

179 },

180 "datetime": {

181 "type": "string",

182 "description": "Reservation date and time (ISO 8601 format)."

183 }

184 },

185 "required": ["name", "datetime"]

186 }

187}

188```

189 

190In the prompt, you may have a section that references the tool like this:

191 

192```text

193<reservation_tool_usage_rules>

194- When the user asks to book, reserve, or schedule a table, you MUST call `create_reservation`.

195- Do NOT guess a reservation time or name — ask for whichever detail is missing.

196- If the user has not provided a name, ask: “What name should I put on the reservation?”

197- If the user has not provided a date/time, ask: “What date and time would you like to reserve?”

198- After calling the tool, confirm the reservation naturally: “Your reservation is confirmed for [name] on [date/time].”

199</tool_usage_rules>

200 

201<reservation_tool_example>

202*Example 1:*

203User: “Book a table for Sarah tomorrow at 7pm.”

204Assistant → (calls tool) →

205`{"name": "create_reservation", "arguments": { "name": "Sarah", "datetime": "2025-11-01T19:00" } }`

206Tool returns: `{ "confirmation_number": "R12345" }`

207Assistant: “All set — your reservation for Sarah tomorrow at 7:00pm is confirmed. Your confirmation number is R12345.”

208 

209**Example 2:**

210User: “I want to make a reservation.”

211Assistant: “Sure! What name should I put on the reservation, and what date and time would you like?”

212 

213*Example 3:*

214User: “Reserve a table under Daniel at 6 tonight.”

215Assistant → (calls tool) →

216`{"name": "create_reservation", "arguments": { "name": "Daniel", "datetime": "2025-10-31T18:00" } }`

217Tool returns: `{ "confirmation_number": "R67890" }`

218Assistant: “Done! Your reservation for Daniel at 6:00pm tonight is confirmed. The confirmation number is R67890.”

219</reservation_tool_example>

220```

221 

222GPT-5.1 also executes parallel tool calls more efficiently. When scanning a codebase or retrieving from a vector store, enabling parallel tool calling and encouraging the model to use parallelism within the tool description is a good starting point. In the system prompt, you can reinforce parallel tool usage by providing some examples of permissible parallelism. An example instruction may look like:

223 

224```text

225Parallelize tool calls whenever possible. Batch reads (read_file) and edits (apply_patch) to speed up the process.

226```

227 

228#### Using the “none” reasoning mode for improved efficiency

229 

230GPT-5.1 introduces a new reasoning mode: `none`. Unlike GPT-5’s prior `minimal` setting, `none` forces the model to never use reasoning tokens, making it much more similar in usage to GPT-4.1, GPT-4o, and other prior non-reasoning models. Importantly, developers can now use hosted tools like [web search](https://developers.openai.com/api/docs/guides/tools-web-search?api-mode=responses) and [file search](https://developers.openai.com/api/docs/guides/tools?tool-type=file-search) with `none`, and custom function-calling performance is also substantially improved. With that in mind, [prior guidance on prompting non-reasoning models](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide) like GPT-4.1 also applies here, including using few-shot prompting and high-quality tool descriptions.

231 

232While GPT-5.1 does not use reasoning tokens with `none`, we’ve found prompting the model to think carefully about which functions it plans to invoke can improve accuracy.

233 

234```text

235You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls, ensuring user's query is completely resolved. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully. In addition, ensure function calls have the correct arguments.

236```

237 

238We’ve also observed that on longer model execution, encouraging the model to “verify” its outputs results in better instruction following for tool use. Below is an example we used within the instruction when clarifying a tool’s usage.

239 

240```text

241When selecting a replacement variant, verify it meets all user constraints (cheapest, brand, spec, etc.). Quote the item-id and price back for confirmation before executing.

242```

243 

244In our testing, GPT-5’s prior `minimal` reasoning mode sometimes led to executions that terminated prematurely. Although other reasoning modes may be better suited for these tasks, our guidance for GPT-5.1 with `none` is similar. Below is a snippet from our Tau bench prompt.

245 

246```text

247Remember, you are an agent - please keep going until the user’s query is completely resolved, before ending your turn and yielding back to the user. You must be prepared to answer multiple queries and only finish the call once the user has confirmed they're done.

248```

249 

250### Maximizing coding performance from planning to execution

251 

252One tool we recommend implementing for long-running tasks is a planning tool. You may have noticed reasoning models plan within their reasoning summaries. Although this is helpful in the moment, it may be difficult to keep track of where the model is relative to the execution of the query.

253 

254```text

255<plan_tool_usage>

256- For medium or larger tasks (e.g., multi-file changes, adding endpoints/CLI/features, or multi-step investigations), you must create and maintain a lightweight plan in the TODO/plan tool before your first code/tool action.

257- Create 2–5 milestone/outcome items; avoid micro-steps and repetitive operational tasks (no “open file”, “run tests”, or similar operational steps). Never use a single catch-all item like “implement the entire feature”.

258- Maintain statuses in the tool: exactly one item in_progress at a time; mark items complete when done; post timely status transitions (never more than ~8 tool calls without an update). Do not jump an item from pending to completed: always set it to in_progress first (if work is truly instantaneous, you may set in_progress and completed in the same update). Do not batch-complete multiple items after the fact.

259- Finish with all items completed or explicitly canceled/deferred before ending the turn.

260- End-of-turn invariant: zero in_progress and zero pending; complete or explicitly cancel/defer anything remaining with a brief reason.

261- If you present a plan in chat for a medium/complex task, mirror it into the tool and reference those items in your updates.

262- For very short, simple tasks (e.g., single-file changes ≲ ~10 lines), you may skip the tool. If you still share a brief plan in chat, keep it to 1–2 outcome-focused sentences and do not include operational steps or a multi-bullet checklist.

263- Pre-flight check: before any non-trivial code change (e.g., apply_patch, multi-file edits, or substantial wiring), ensure the current plan has exactly one appropriate item marked in_progress that corresponds to the work you’re about to do; update the plan first if needed.

264- Scope pivots: if understanding changes (split/merge/reorder items), update the plan before continuing. Do not let the plan go stale while coding.

265- Never have more than one item in_progress; if that occurs, immediately correct the statuses so only the current phase is in_progress.

266<plan_tool_usage>

267```

268 

269A plan tool can be used with minimal scaffolding. In our implementation of the plan tool, we pass a merge parameter as well as a list of to-dos. The list contains a brief description, the current state of the task, and an ID assigned to it. Below is an example of a function call that GPT-5.1 may make to record its state.

270 

271```json

272{

273 "name": "update_plan",

274 "arguments": {

275 "merge": true,

276 "todos": [

277 {

278 "content": "Investigate failing test",

279 "status": "in_progress",

280 "id": "step-1"

281 },

282 {

283 "content": "Apply fix and re-run tests",

284 "status": "pending",

285 "id": "step-2"

286 }

287 ]

288 }

289}

290```

291 

292#### Design system enforcement

293 

294When building frontend interfaces, GPT-5.1 can be steered to produce websites that match your visual design system. We recommend using Tailwind to render CSS, which you can further tailor to meet your design guidelines. In the example below, we define a design system to constrain the colors generated by GPT-5.1.

295 

296```text

297<design_system_enforcement>

298- Tokens-first: Do not hard-code colors (hex/hsl/oklch/rgb) in JSX/CSS. All colors must come from globals.css variables (e.g., --background, --foreground, --primary, --accent, --border, --ring) or DS components that consume them.

299- Introducing a brand or accent? Before styling, add/extend tokens in globals.css under :root and .dark, for example:

300 - --brand, --brand-foreground, optional --brand-muted, --brand-ring, --brand-surface

301 - If gradients/glows are needed, define --gradient-1, --gradient-2, etc., and ensure they reference sanctioned hues.

302- Consumption: Use Tailwind/CSS utilities wired to tokens (e.g., bg-[hsl(var(--primary))], text-[hsl(var(--foreground))], ring-[hsl(var(--ring))]). Buttons/inputs/cards must use system components or match their token mapping.

303- Default to the system's neutral palette unless the user explicitly requests a brand look; then map that brand to tokens first.

304</design_system_enforcement>

305```

306 

307### New tool types in GPT-5.1

308 

309GPT-5.1 has been post-trained on specific tools that are commonly used in coding use cases. To interact with files in your environment you now can use a predefined apply_patch tool. Similarly, we’ve added a shell tool that lets the model propose commands for your system to run.

310 

311#### Using apply_patch

312 

313The apply_patch tool lets GPT-5.1 create, update, and delete files in your codebase using structured diffs. Instead of just suggesting edits, the model emits patch operations that your application applies and then reports back on, enabling iterative, multi-step code editing workflows. You can find additional usage details and context in the [GPT-4.1 prompting guide](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide#:~:text=PYTHON_TOOL_DESCRIPTION%20%3D%20%22%22%22This,an%20exclamation%20mark.).

314 

315With GPT-5.1, you can use apply_patch as a new tool type without writing custom descriptions for the tool. The description and handling are managed via the Responses API. Under the hood, this implementation uses a freeform function call rather than a JSON format. In testing, the named function decreased apply_patch failure rates by 35%.

316 

317```python

318response = client.responses.create(

319 model="gpt-5.1", input=RESPONSE_INPUT, tools=[{"type": "apply_patch"}]

320)

321```

322 

323```java

324import com.openai.client.OpenAIClient;

325import com.openai.client.okhttp.OpenAIOkHttpClient;

326import com.openai.models.responses.ApplyPatchTool;

327import com.openai.models.responses.ResponseCreateParams;

328 

329ResponseCreateParams params =

330 ResponseCreateParams.builder()

331 .model("gpt-5.1")

332 .input("Update the README title and fix the failing test.")

333 .addTool(ApplyPatchTool.builder().build())

334 .build();

335 

336client.responses().create(params).output().stream()

337 .flatMap(item -> item.message().stream())

338 .flatMap(message -> message.content().stream())

339 .flatMap(content -> content.outputText().stream())

340 .forEach(text -> System.out.println(text.text()));

341```

342 

343```ruby

344response = client.responses.create(

345 model: "gpt-5.1", input: response_input, tools: [{ type: :apply_patch }]

346)

347```

348 

349 

350When the model decides to execute an apply_patch tool, you will receive an apply_patch_call function type within the response stream. Within the operation object, you’ll receive a type field (with one of `create_file`, `update_file`, or `delete_file`) and the diff to implement.

351 

352```text

353{

354 "id": "apc_08f3d96c87a585390069118b594f7481a088b16cda7d9415fe",

355 "type": "apply_patch_call",

356 "status": "completed",

357 "call_id": "call_Rjsqzz96C5xzPb0jUWJFRTNW",

358 "operation": {

359 "type": "update_file",

360 "diff": "

361 @@

362 -def fib(n):

363 +def fibonacci(n):

364 if n <= 1:

365 return n

366 - return fib(n-1) + fib(n-2)

367 + return fibonacci(n-1) + fibonacci(n-2)",

368 "path": "lib/fib.py"

369 }

370},

371 

372```

373 

374[This repository](https://github.com/openai/openai-cookbook/blob/main/examples/gpt-5/apply_patch.py) contains the expected implementation for the apply_patch tool executable. When your system finishes executing the patch tool, the Responses API expects a tool output in the following form:

375 

376```python

377{

378 "type": "apply_patch_call_output",

379 "call_id": call["call_id"],

380 "status": "completed" if success else "failed",

381 "output": log_output,

382}

383```

384 

385```ruby

386output = {

387 type: :apply_patch_call_output,

388 call_id: call_id,

389 status: success ? :completed : :failed,

390 output: log_output

391}

392```

393 

394 

395#### Using the shell tool

396 

397We’ve also built a new shell tool for GPT-5.1. The shell tool allows the model to interact with your local computer through a controlled command-line interface. The model proposes shell commands; your integration executes them and returns the outputs. This creates a simple plan-execute loop that lets models inspect the system, run utilities, and gather data until they finish the task.

398 

399The shell tool is invoked in the same way as apply_patch: include it as a tool of type `shell`.

400 

401```python

402tools = [{"type": "shell"}]

403```

404 

405```ruby

406tools = [{ type: :shell }]

407```

408 

409 

410When a shell tool call is returned, the Responses API includes a `shell_call` object with a timeout, a maximum output length, and the command to run.

411 

412```text

413{

414 "type": "shell_call",

415 "call_id": "...",

416 "action": {

417 "commands": [...],

418 "timeout_ms": 120000,

419 "max_output_length": 4096

420 },

421 "status": "in_progress"

422}

423```

424 

425After executing the shell command, return the untruncated stdout/stderr logs as well as the exit-code details.

426 

427```json

428{

429 "type": "shell_call_output",

430 "call_id": "...",

431 "max_output_length": 4096,

432 "output": [

433 {

434 "stdout": "...",

435 "stderr": "...",

436 "outcome": {

437 "type": "exit",

438 "exit_code": 0

439 }

440 }

441 ]

442}

443```

444 

445### How to metaprompt effectively

446 

447Building prompts can be cumbersome, but it’s also the highest-leverage thing you can do to resolve most model behavior issues. Small inclusions can unexpectedly steer the model undesirably. Let’s walk through an example of an agent that plans events. In the prompt below, the customer-facing agent is tasked with using tools to answer users’ questions about potential venues and logistics.

448 

449```text

450You are “GreenGather,” an autonomous sustainable event-planning agent. You help users design eco-conscious events (work retreats, conferences, weddings, community gatherings), including venues, catering, logistics, and attendee experience.

451 

452PRIMARY OBJECTIVE

453Your main goal is to produce concise, immediately actionable answers that fit in a quick chat context. Most responses should be about 3–6 sentences total. Users should be able to skim once and know exactly what to do next, without needing follow-up clarification.

454 

455SCOPE

456 

457* Focus on: venue selection, schedule design, catering styles, transportation choices, simple budgeting, and sustainability considerations.

458* You do not actually book venues or vendors; never say you completed a booking.

459* You may, however, phrase suggestions as if the user can follow them directly (“Book X, then do Y”) so planning feels concrete and low-friction.

460 

461TONE & STYLE

462 

463* Sound calm, professional, and neutral, suitable for corporate planners and executives. Avoid emojis and expressive punctuation.

464* Do not use first-person singular; prefer “A good option is…” or “It is recommended that…”.

465* Be warm and approachable. For informal or celebratory events (e.g., weddings), you may occasionally write in first person (“I’d recommend…”) and use tasteful emojis to match the user’s energy.

466 

467STRUCTURE

468Default formatting guidelines:

469 

470* Prefer short paragraphs, not bullet lists.

471* Use bullets only when the user explicitly asks for “options,” “list,” or “checklist.”

472* For complex, multi-day events, always structure your answer with labeled sections (e.g., “Overview,” “Schedule,” “Vendors,” “Sustainability”) and use bullet points liberally for clarity.

473 

474AUTONOMY & PLANNING

475You are an autonomous agent. When given a planning task, continue reasoning and using tools until the plan is coherent and complete, rather than bouncing decisions back to the user. Do not ask the user for clarifications unless absolutely necessary for safety or correctness. Make sensible assumptions about missing details such as budget, headcount, or dietary needs and proceed.

476 

477To avoid incorrect assumptions, when key information (date, city, approximate headcount) is missing, pause and ask 1–3 brief clarifying questions before generating a detailed plan. Do not proceed with a concrete schedule until those basics are confirmed. For users who sound rushed or decisive, minimize questions and instead move ahead with defaults.

478 

479TOOL USAGE

480You always have access to tools for:

481 

482* venue_search: find venues with capacity, location, and sustainability tags

483* catering_search: find caterers and menu styles

484* transport_search: find transit and shuttle options

485* budget_estimator: estimate costs by category

486 

487General rules for tools:

488 

489* Prefer tools over internal knowledge whenever you mention specific venues, vendors, or prices.

490* For simple conceptual questions (e.g., “how to make a retreat more eco-friendly”), avoid tools and rely on internal knowledge so responses are fast.

491* For any event with more than 30 attendees, always call at least one search tool to ground recommendations in realistic options.

492* To keep the experience responsive, avoid unnecessary tool calls; for rough plans or early brainstorming, you can freely propose plausible example venues or caterers from general knowledge instead of hitting tools.

493 

494When using tools as an autonomous agent:

495 

496* Plan your approach (which tools, in what order) and then execute without waiting for user confirmation at each step.

497* After each major tool call, briefly summarize what you did and how results shaped your recommendation.

498* Keep tool usage invisible unless the user explicitly asks how you arrived at a suggestion.

499 

500VERBOSITY & DETAIL

501Err on the side of completeness so the user does not need follow-up messages. Include specific examples (e.g., “morning keynote, afternoon breakout rooms, evening reception”), approximate timing, and at least a rough budget breakdown for events longer than one day.

502 

503However, respect the user’s time: long walls of text are discouraged. Aim for compact responses that rarely exceed 2–3 short sections. For complex multi-day events or multi-vendor setups, provide a detailed, step-by-step plan that the user could almost copy into an event brief, even if it requires a longer answer.

504 

505SUSTAINABILITY GUIDANCE

506 

507* Whenever you suggest venues or transportation, include at least one lower-impact alternative (e.g., public transit, shuttle consolidation, local suppliers).

508* Do not guilt or moralize; frame tradeoffs as practical choices.

509* Highlight sustainability certifications when relevant, but avoid claiming a venue has a certification unless you are confident based on tool results or internal knowledge.

510 

511INTERACTION & CLOSING

512Avoid over-apologizing or repeating yourself. Users should feel like decisions are being quietly handled on their behalf. Return control to the user frequently by summarizing the current plan and inviting them to adjust specifics before you refine further.

513 

514End every response with a subtle next step the user could take, phrased as a suggestion rather than a question, and avoid explicit calls for confirmation such as “Let me know if this works.”

515```

516 

517Although this is a strong starting prompt, there are a few issues we noticed upon testing:

518 

519- Small conceptual questions (like asking about a 20-person leadership dinner) triggered unnecessary tool calls and very concrete venue suggestions, despite the prompt allowing internal knowledge for simple, high-level questions.

520 

521- The agent oscillated between being overly verbose (multi-day Austin offsites turning into dense, multi-section essays) and overly hesitant (refusing to propose a plan without more questions) and occasionally ignored unit rules (a Berlin summit described in miles and °F instead of km and °C).

522 

523Rather than manually guessing which lines of the system prompt caused these behaviors, we can metaprompt GPT-5.1 to inspect its own instructions and traces.

524 

525**Step 1**: Ask GPT-5.1 to diagnose failures

526 

527Paste the system prompt and a small batch of failure examples into a separate analysis call. Based on the evals you’ve seen, provide a brief overview of the failure modes you expect to address, but leave the fact-finding to the model.

528 

529Note that in this prompt, we’re not asking for a solution yet, just a root-cause analysis.

530 

531```text

532You are a prompt engineer tasked with debugging a system prompt for an event-planning agent that uses tools to recommend venues, logistics, and sustainable options.

533 

534You are given:

535 

5361) The current system prompt:

537<system_prompt>

538[DUMP_SYSTEM_PROMPT]

539</system_prompt>

540 

5412) A small set of logged failures. Each log has:

542- query

543- tools_called (as actually executed)

544- final_answer (shortened if needed)

545- eval_signal (e.g., thumbs_down, low rating, human grader, or user comment)

546 

547<failure_tracess>

548[DUMP_FAILURE_TRACES]

549</failure_traces>

550 

551Your tasks:

552 

5531) Identify the distinct failure mode you see (e.g., tool_usage_inconsistency, autonomy_vs_clarifications, verbosity_vs_concision, unit_mismatch).

5542) For each failure mode, quote or paraphrase the specific lines or sections of the system prompt that are most likely causing or reinforcing it. Include any contradictions (e.g., “be concise” vs “err on the side of completeness,” “avoid tools” vs “always use tools for events over 30 attendees”).

5553) Briefly explain, for each failure mode, how those lines are steering the agent toward the observed behavior.

556 

557Return your answer in a structured but readable format:

558 

559failure_modes:

560- name: ...

561 description: ...

562 prompt_drivers:

563 - exact_or_paraphrased_line: ...

564 - why_it_matters: ...

565```

566 

567Metaprompting works best when the feedback can logically be grouped together. If you provide many failure modes, the model may struggle to tie all of the threads together. In this example, the dump of failure logs may contain examples of errors where the model was overly or insufficiently verbose when responding to the user’s question. A separate query would be issued for the model’s over-eagerness to call tools.

568 

569**Step 2:** Ask GPT-5.1 how it would patch the prompt to fix those behaviors

570 

571Once you have that analysis, you can run a second, separate call that focuses on implementation: tightening the prompt without fully rewriting it.

572 

573```text

574You previously analyzed this system prompt and its failure modes.

575 

576System prompt:

577 

578<system_prompt>

579[DUMP_SYSTEM_PROMPT]

580</system_prompt>

581 

582Failure-mode analysis:

583[DUMP_FAILURE_MODE_ANALYSIS]

584 

585Please propose a surgical revision of the system prompt that reduces the observed issues while preserving the good behaviors.

586 

587Constraints:

588 

589- Do not redesign the agent from scratch.

590- Prefer small, explicit edits: clarify conflicting rules, remove redundant or contradictory lines, tighten vague guidance.

591- Make tradeoffs explicit (for example, clearly state when to prioritize concision over completeness, or exactly when tools must vs must not be called).

592- Keep the structure and overall length roughly similar to the original, unless a short consolidation removes obvious duplication.

593 

594Output:

595 

5961) patch_notes: a concise list of the key changes and the reasoning behind each (e.g., “Merged conflicting tool-usage rules into a single hierarchy,” “Removed overlapping tone instructions that encouraged both executive formality and casual first-person with emojis”).

5972) revised_system_prompt: the full updated system prompt with your edits applied, ready to drop into an agent configuration.

598```

599 

600In this example, the first metaprompt helps GPT-5.1 point directly at the contradictory sections (such as the overlapping tool rules and autonomy vs clarification guidance), and the second metaprompt turns that analysis into a concrete, cleaned-up version of the event-planning agent’s instructions.

601 

602The output from the second prompt might look something like this:

603 

604```text

605patch_notes:

606- Clarified when to prioritize concision vs detail:

607 - Simple or single-topic queries should stay within ~3–6 sentences.

608 - Longer, structured answers are reserved for clearly complex, multi-day or multi-vendor requests.

609- Removed language that told the agent to “err on the side of completeness” for all cases and replaced it with conditional rules based on query complexity.

610- Tightened the structure rules so headings and bullets are only used when complexity or explicit user instructions justify them.

611- Simplified the guidance on step-by-step plans so they are expected only for complex events, not for every question.

612 

613revised_system_prompt:

614[...]

615```

616 

617After this iteration cycle, run the queries again to observe any regressions and repeat this process until your failure modes have been identified and triaged.

618 

619As you continue to grow your agentic systems (e.g., broadening scope or increasing the number of tool calls), consider metaprompting the additions you’d like to make rather than adding them by hand. This helps maintain discrete boundaries for each tool and when they should be used.

620 

621### What's next

622 

623To summarize, GPT-5.1 builds on the foundation set by GPT-5 and adds things like quicker thinking for easy questions, steerability when it comes to model output, new tools for coding use cases, and the option to set reasoning to `none` when your tasks don't require heavy thinking.

624 

625Review the [GPT-5.1 model and API guidance](#model-api-and-feature-updates), or read the [blog post](https://openai.com/index/gpt-5-1-for-developers/) to learn more.

626 

guides/latest-model/gpt-5.2.md +0 −1079 deleted

File Deleted View Diff

1# Using GPT-5.2

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7GPT-5.2 was released as a flagship general-purpose model for both general and agentic tasks. Compared with GPT-5.1, it improved:

8 

9- General intelligence

10- Instruction following

11- Accuracy and token efficiency

12- Multimodality—especially vision

13- Code generation—especially front-end UI creation

14- Tool calling and context management in the API

15- Spreadsheet understanding and creation

16 

17Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers" to improve accuracy.

18 

19This guide covers key features of the GPT-5 model family and how to get the most out of GPT-5.2.

20 

21## Explore coding examples

22 

23Click through a few demo applications generated entirely with a single prompt, without writing any code by hand. Note that these examples were either generated by GPT-5.2 or our previous flagship model, GPT-5.

24 

25## Model, API, and feature updates

26 

27The GPT-5.2 generation includes `gpt-5.2` for complex tasks that require broad world knowledge, `gpt-5.2-chat-latest` for ChatGPT-aligned behavior, and `gpt-5.2-pro` for problems that benefit from more compute.

28 

29For a smaller model, use `gpt-5-mini`.

30 

31To help you pick the model that best fits your use case, consider these tradeoffs:

32 

33| Variant | Best for |

34| ------------------------------------------------- | ------------------------------------------------------------------------------------ |

35| [`gpt-5.2`](https://developers.openai.com/api/docs/models/gpt-5.2) | Complex reasoning, broad world knowledge, and code-heavy or multi-step agentic tasks |

36| [`gpt-5.2-pro`](https://developers.openai.com/api/docs/models/gpt-5.2-pro) | Tough problems that may take longer to solve but require harder thinking |

37| [`gpt-5.2-codex`](https://developers.openai.com/api/docs/models/gpt-5.2-codex) | Companies building interactive coding products; full spectrum of coding tasks |

38| [`gpt-5-mini`](https://developers.openai.com/api/docs/models/gpt-5-mini) | Cost-optimized reasoning and chat; balances speed, cost, and capability |

39| [`gpt-5-nano`](https://developers.openai.com/api/docs/models/gpt-5-nano) | High-throughput tasks, especially focused instruction-following or classification |

40 

41### New features in GPT-5.2

42 

43Just like GPT-5.1, the new GPT-5.2 has API features like custom tools, parameters to control verbosity and reasoning, and an allowed tools list. What's new in 5.2 is a new `xhigh` reasoning effort level, concise reasoning summaries, and new context management using _compaction_.

44 

45This guide walks through some of the key features of the GPT-5 model family and how to get the most out of 5.2 in particular.

46 

47For coding tasks, GPT-5.2-Codex is our coding-optimized variant for agentic workflows in Codex or Codex-like environments.

48 

49### Lower reasoning effort

50 

51The `reasoning.effort` parameter controls how many reasoning tokens the model generates before producing a response. Earlier reasoning models like o3 supported only `low`, `medium`, and `high`: `low` favored speed and fewer tokens, while `high` favored more thorough reasoning.

52 

53With GPT-5.2, the lowest setting is `none` to provide lower-latency interactions. This is the default setting in GPT-5.2. If you need more thinking, slowly increase to `medium` and experiment with results.

54 

55With reasoning effort set to `none`, prompting is important. To improve the model's reasoning quality, even with the default settings, encourage it to “think” or outline its steps before answering.

56 

57Reasoning effort set to none

58 

59```javascript

60import OpenAI from "openai";

61const openai = new OpenAI();

62 

63const response = await openai.responses.create({

64 model: "gpt-5.2",

65 input:

66 "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

67 reasoning: {

68 effort: "none",

69 },

70});

71 

72console.log(response);

73```

74 

75```python

76from openai import OpenAI

77 

78client = OpenAI()

79 

80response = client.responses.create(

81 model="gpt-5.2",

82 input="Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

83 reasoning={"effort": "none"},

84)

85 

86print(response)

87```

88 

89```go

90package main

91 

92import (

93 "context"

94 "fmt"

95 

96 "github.com/openai/openai-go/v3"

97 "github.com/openai/openai-go/v3/responses"

98 "github.com/openai/openai-go/v3/shared"

99)

100 

101func main() {

102 client := openai.NewClient()

103 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{

104 Model: "gpt-5.2",

105 Input: responses.ResponseNewParamsInputUnion{OfString: openai.String("Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?")},

106 Reasoning: shared.ReasoningParam{Effort: shared.ReasoningEffortNone},

107 })

108 if err != nil {

109 panic(err)

110 }

111 fmt.Println(response)

112}

113```

114 

115```java

116import com.openai.client.OpenAIClient;

117import com.openai.client.okhttp.OpenAIOkHttpClient;

118import com.openai.models.Reasoning;

119import com.openai.models.ReasoningEffort;

120import com.openai.models.responses.ResponseCreateParams;

121 

122ResponseCreateParams params =

123 ResponseCreateParams.builder()

124 .model("gpt-5.2")

125 .input("Explain the bug and propose a fix.")

126 .reasoning(Reasoning.builder().effort(ReasoningEffort.NONE).build())

127 .build();

128 

129client.responses().create(params).output().stream()

130 .flatMap(item -> item.message().stream())

131 .flatMap(message -> message.content().stream())

132 .flatMap(content -> content.outputText().stream())

133 .forEach(text -> System.out.println(text.text()));

134```

135 

136```csharp

137using OpenAI.Responses;

138#pragma warning disable OPENAI001

139 

140string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;

141ResponsesClient client = new(key);

142 

143CreateResponseOptions options = new()

144{

145 Model = "gpt-5.2",

146 ReasoningOptions = new ResponseReasoningOptions

147 {

148 ReasoningEffortLevel = ResponseReasoningEffortLevel.None,

149 },

150};

151options.InputItems.Add(

152 ResponseItem.CreateUserMessageItem(

153 "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?"

154 )

155);

156 

157ResponseResult response = await client.CreateResponseAsync(options);

158Console.WriteLine(response.GetOutputText());

159```

160 

161```ruby

162require "openai"

163 

164client = OpenAI::Client.new

165response = client.responses.create(

166 model: "gpt-5.2",

167 reasoning: { effort: :minimal },

168 input: "Explain the bug and propose a fix."

169)

170puts(response.output_text)

171```

172 

173```bash

174curl --request POST \

175 --url https://api.openai.com/v1/responses \

176 --header "Authorization: Bearer $OPENAI_API_KEY" \

177 --header 'Content-type: application/json' \

178 --data '{

179 "model": "gpt-5.2",

180 "input": "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

181 "reasoning": {

182 "effort": "none"

183 }

184}'

185```

186 

187 

188### Verbosity

189 

190Verbosity determines how many output tokens are generated. Lowering the number of tokens reduces overall latency. While the model's reasoning approach stays mostly the same, the model finds ways to answer more concisely—which can either improve or diminish answer quality, depending on your use case. Here are some scenarios for both ends of the verbosity spectrum:

191 

192- **High verbosity:** Use when you need the model to provide thorough explanations of documents or perform extensive code refactoring.

193- **Low verbosity:** Best for situations where you want concise answers or focused code generation, such as SQL queries.

194 

195GPT-5 made this option configurable as one of `high`, `medium`, or `low`. With GPT-5.2, verbosity remains configurable and defaults to `medium`.

196 

197When generating code with GPT-5.2, `medium` and `high` verbosity levels yield longer, more structured code with inline explanations, while `low` verbosity produces shorter, more concise code with minimal commentary.

198 

199Control verbosity

200 

201```javascript

202import OpenAI from "openai";

203const openai = new OpenAI();

204 

205const response = await openai.responses.create({

206 model: "gpt-5.2",

207 input:

208 "What is the answer to the ultimate question of life, the universe, and everything?",

209 text: {

210 verbosity: "low",

211 },

212});

213 

214console.log(response);

215```

216 

217```python

218from openai import OpenAI

219 

220client = OpenAI()

221 

222response = client.responses.create(

223 model="gpt-5.2",

224 input="What is the answer to the ultimate question of life, the universe, and everything?",

225 text={"verbosity": "low"},

226)

227 

228print(response)

229```

230 

231```go

232package main

233 

234import (

235 "context"

236 "fmt"

237 

238 "github.com/openai/openai-go/v3"

239 "github.com/openai/openai-go/v3/responses"

240)

241 

242func main() {

243 client := openai.NewClient()

244 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{

245 Model: "gpt-5.2",

246 Input: responses.ResponseNewParamsInputUnion{OfString: openai.String("What is the answer to the ultimate question of life, the universe, and everything?")},

247 Text: responses.ResponseTextConfigParam{Verbosity: responses.ResponseTextConfigVerbosityLow},

248 })

249 if err != nil {

250 panic(err)

251 }

252 fmt.Println(response)

253}

254```

255 

256```java

257import com.openai.client.OpenAIClient;

258import com.openai.client.okhttp.OpenAIOkHttpClient;

259import com.openai.models.responses.ResponseCreateParams;

260import com.openai.models.responses.ResponseTextConfig;

261 

262ResponseCreateParams params =

263 ResponseCreateParams.builder()

264 .model("gpt-5.2")

265 .input("Explain the bug and propose a fix.")

266 .text(ResponseTextConfig.builder().verbosity(ResponseTextConfig.Verbosity.LOW).build())

267 .build();

268 

269client.responses().create(params).output().stream()

270 .flatMap(item -> item.message().stream())

271 .flatMap(message -> message.content().stream())

272 .flatMap(content -> content.outputText().stream())

273 .forEach(text -> System.out.println(text.text()));

274```

275 

276```ruby

277require "openai"

278 

279client = OpenAI::Client.new

280response = client.responses.create(

281 model: "gpt-5.2",

282 text: { verbosity: :low },

283 input: "Explain the bug and propose a fix."

284)

285puts(response.output_text)

286```

287 

288```bash

289curl --request POST \

290 --url https://api.openai.com/v1/responses \

291 --header "Authorization: Bearer $OPENAI_API_KEY" \

292 --header 'Content-type: application/json' \

293 --data '{

294 "model": "gpt-5.2",

295 "input": "What is the answer to the ultimate question of life, the universe, and everything?",

296 "text": {

297 "verbosity": "low"

298 }

299}'

300```

301 

302 

303You can still steer verbosity through prompting after setting it to `low` in the API. The verbosity parameter defines a general token range at the system prompt level, but the actual output is flexible to both developer and user prompts within that range.

304 

305### Using tools with GPT-5.2

306 

307GPT-5.2 has been post-trained on specific tools. See the [tools docs](https://developers.openai.com/api/docs/guides/tools) for more specific guidance.

308 

309#### The apply patch tool

310 

311The `apply_patch` tool lets GPT-5.2 create, update, and delete files in your codebase using structured diffs. Instead of just suggesting edits, the model emits patch operations that your application applies and then reports back on, enabling iterative, multi-step code editing workflows. [Read the docs](https://developers.openai.com/api/docs/guides/tools-apply-patch).

312 

313Under the hood, this implementation uses a freeform function call rather than a JSON format. In testing, the named function decreased `apply_patch` failure rates by 35%.

314 

315#### Shell tool

316 

317Local shell is supported in GPT-5.2. The shell tool allows the model to interact with your local computer through a controlled command-line interface. [Read the docs](https://developers.openai.com/api/docs/guides/tools-shell) to learn more.

318 

319### Custom tools

320 

321When the GPT-5 model family launched, we introduced a new capability called custom tools, which lets models send any raw text as tool call input but still constrain outputs if desired. This tool behavior remains true in GPT-5.2.

322 

323[Function calling guide

324 

325 

326 

327 Learn about custom tools in the function calling guide.](https://developers.openai.com/api/docs/guides/function-calling)

328 

329#### Freeform inputs

330 

331Define your tool with `type: custom` to enable models to send plaintext inputs directly to your tools, rather than being limited to structured JSON. The model can send any raw text—code, SQL queries, shell commands, configuration files, or long-form prose—directly to your tool.

332 

333```json

334{

335 "type": "custom",

336 "name": "code_exec",

337 "description": "Executes arbitrary python code"

338}

339```

340 

341#### Constraining outputs

342 

343GPT-5.2 supports context-free grammars (`CFGs`) for custom tools, letting you provide a Lark grammar to constrain outputs to a specific syntax or DSL. Attaching a CFG, for example a SQL or DSL grammar, ensures the assistant's text matches your grammar.

344 

345This enables precise, constrained tool calls or structured responses and lets you enforce strict syntactic or domain-specific formats directly in GPT-5.2's function calling, improving control and reliability for complex or constrained domains.

346 

347#### Best practices for custom tools

348 

349- **Write concise, explicit tool descriptions.** The model chooses what to send based on your description; state explicitly if you want it to always call the tool.

350- **Validate outputs on the server side**. Freeform strings are powerful but require safeguards against injection or unsafe commands.

351 

352### Allowed tools

353 

354The `allowed_tools` parameter under `tool_choice` lets you pass N tool definitions but restrict the model to only M (&lt; N) of them. List your full toolkit in `tools`, and then use an `allowed_tools` block to name the subset and specify a mode—either `auto` (the model may pick any of those) or `required` (the model must invoke one).

355 

356[Function calling guide

357 

358 

359 

360 Learn about the allowed tools option in the function calling guide.](https://developers.openai.com/api/docs/guides/function-calling)

361 

362By separating all possible tools from the subset that can be used _now_, you gain greater safety, predictability, and improved prompt caching. You also avoid brittle prompt engineering, such as hard-coded call order. GPT-5.2 dynamically invokes or requires specific functions mid-conversation while reducing the risk of unintended tool usage over long contexts.

363 

364| | **Standard Tools** | **Allowed Tools** |

365| ---------------- | ----------------------------------------- | ------------------------------------------------------------- |

366| Model's universe | All tools listed under **`"tools": […]`** | Only the subset under **`"tools": […]`** in **`tool_choice`** |

367| Tool invocation | Model may or may not call any tool | Model restricted to (or required to call) chosen tools |

368| Purpose | Declare available capabilities | Constrain which capabilities are actually used |

369 

370```json

371{

372 "tool_choice": {

373 "type": "allowed_tools",

374 "mode": "auto",

375 "tools": [

376 { "type": "function", "name": "get_weather" },

377 { "type": "function", "name": "search_docs" }

378 ]

379 }

380}

381```

382 

383For a more detailed overview of all of these new features, see the [accompanying cookbook](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5-2_prompting_guide).

384 

385### Preambles

386 

387Preambles are brief, user-visible explanations that GPT-5.2 generates before invoking any tool or function, outlining its intent or plan—for example, “why I'm calling this tool.” They appear after the chain of thought and before the actual tool call, making the model's reasoning easier to understand and debug while supporting precise steering.

388 

389By letting GPT-5.2 “think out loud” before each tool call, preambles boost tool-calling accuracy (and overall task success) without bloating reasoning overhead. To enable preambles, add a system or developer instruction—for example: “Before you call a tool, explain why you are calling it.” GPT-5.2 adds a concise rationale to each specified tool call. The model may also output multiple messages between tool calls, which can enhance the interaction experience—particularly for minimal reasoning or latency-sensitive use cases.

390 

391For more on using preambles, see the [GPT-5 prompting cookbook](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_prompting_guide#tool-preambles).

392 

393## Migration quickstart

394 

395GPT-5.2 works best with the Responses API, which supports preserving reasoning context between turns. Read below to migrate from your current model or API.

396 

397### Migrating from other models to GPT-5.2

398 

399While the model should be close to a drop-in replacement for GPT-5.1, there are a few key changes to call out. See the [GPT-5.2 prompting guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5-2_prompting_guide) for specific updates to make in your prompts.

400 

401Using GPT-5 models with the Responses API provides improved intelligence because of the API design. The Responses API can pass the previous turn's CoT to the model. This leads to fewer generated reasoning tokens, higher cache hit rates, and less latency. To learn more, see an [in-depth guide](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items) on the benefits of the Responses API.

402 

403When migrating to GPT-5.2 from an older OpenAI model, start by experimenting with reasoning levels and prompting strategies. Based on our testing, we recommend using our [prompt optimizer](https://platform.openai.com/chat/edit?models=gpt-5.2&optimize=true)—which automatically updates your prompts for GPT-5.2 based on our best practices—and following this model-specific guidance:

404 

405- **`gpt-5.1`**: `gpt-5.2` with default settings is meant to be a drop-in replacement.

406- **o3**: `gpt-5.2` with `medium` or `high` reasoning. Start with `medium` reasoning with prompt tuning, then increase to `high` if you aren't getting the results you want.

407- **`gpt-4.1`**: `gpt-5.2` with `none` reasoning. Start with `none` and tune your prompts; increase if you need better performance.

408- **`o4-mini` or `gpt-4.1-mini`**: `gpt-5-mini` with prompt tuning is a great replacement.

409- **`gpt-4.1-nano`**: `gpt-5-nano` with prompt tuning is a great replacement.

410 

411### GPT-5.2 parameter compatibility

412 

413The following parameters are **only supported** when using GPT-5.2 with reasoning effort set to `none`:

414 

415- `temperature`

416- `top_p`

417- `logprobs`

418 

419Requests to GPT-5.2 or GPT-5.1 with any other reasoning effort setting, or to older GPT-5 models—for example, `gpt-5`, `gpt-5-mini`, or `gpt-5-nano`—that include these fields will raise an error.

420 

421To achieve similar results with reasoning effort set higher, or with another GPT-5 family model, try these alternative parameters:

422 

423- **Reasoning depth:** `reasoning: { effort: "none" | "low" | "medium" | "high" | "xhigh" }`

424- **Output verbosity:** `text: { verbosity: "low" | "medium" | "high" }`

425- **Output length:** `max_output_tokens`

426 

427### Migrating from Chat Completions to Responses API

428 

429The biggest difference, and main reason to migrate from Chat Completions to the Responses API for GPT-5.2, is support for passing chain of thought (CoT) between turns. See a full [comparison of the APIs](https://developers.openai.com/api/docs/guides/migrate-to-responses).

430 

431Passing CoT exists only in the Responses API, and we've seen improved intelligence, fewer generated reasoning tokens, higher cache hit rates, and lower latency as a result of doing so. Most other parameters remain at parity, though the formatting is different. Here's how new parameters are handled differently between Chat Completions and the Responses API:

432 

433**Reasoning effort**

434 

435 

436 

437Responses API

438 

439 Generate response with reasoning effort set to none

440 

441```bash

442curl --request POST \

443 --url https://api.openai.com/v1/responses \

444 --header "Authorization: Bearer $OPENAI_API_KEY" \

445 --header "Content-type: application/json" \

446 --data '{

447 "model": "gpt-5.2",

448 "input": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

449 "reasoning": {

450 "effort": "none"

451 }

452}'

453```

454 

455

456 

457

458 

459

460Chat Completions

461 

462 Generate response with reasoning effort set to none

463 

464```bash

465curl --request POST \

466 --url https://api.openai.com/v1/chat/completions \

467 --header "Authorization: Bearer $OPENAI_API_KEY" \

468 --header "Content-type: application/json" \

469 --data '{

470 "model": "gpt-5.2",

471 "messages": [

472 {

473 "role": "user",

474 "content": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?"

475 }

476 ],

477 "reasoning_effort": "none"

478}'

479```

480 

481 

482 

483**Verbosity**

484 

485 

486 

487Responses API

488 

489 Control verbosity

490 

491```bash

492curl --request POST \

493 --url https://api.openai.com/v1/responses \

494 --header "Authorization: Bearer $OPENAI_API_KEY" \

495 --header "Content-type: application/json" \

496 --data '{

497 "model": "gpt-5.2",

498 "input": "What is the answer to the ultimate question of life, the universe, and everything?",

499 "text": {

500 "verbosity": "low"

501 }

502}'

503```

504 

505

506 

507

508 

509

510Chat Completions

511 

512 Control verbosity

513 

514```bash

515curl --request POST \

516 --url https://api.openai.com/v1/chat/completions \

517 --header "Authorization: Bearer $OPENAI_API_KEY" \

518 --header "Content-type: application/json" \

519 --data '{

520 "model": "gpt-5.2",

521 "messages": [

522 {

523 "role": "user",

524 "content": "What is the answer to the ultimate question of life, the universe, and everything?"

525 }

526 ],

527 "verbosity": "low"

528}'

529```

530 

531 

532 

533**Custom tools**

534 

535 

536 

537Responses API

538 

539 Custom tool call

540 

541```bash

542curl --request POST \

543 --url https://api.openai.com/v1/responses \

544 --header "Authorization: Bearer $OPENAI_API_KEY" \

545 --header "Content-type: application/json" \

546 --data '{

547 "model": "gpt-5.2",

548 "input": "Use the code_exec tool to calculate the area of a circle with radius equal to the number of r letters in blueberry",

549 "tools": [

550 {

551 "type": "custom",

552 "name": "code_exec",

553 "description": "Executes arbitrary Python code"

554 }

555 ]

556}'

557```

558 

559

560 

561

562 

563

564Chat Completions

565 

566 Custom tool call

567 

568```bash

569curl --request POST \

570 --url https://api.openai.com/v1/chat/completions \

571 --header "Authorization: Bearer $OPENAI_API_KEY" \

572 --header "Content-type: application/json" \

573 --data '{

574 "model": "gpt-5.2",

575 "messages": [

576 {

577 "role": "user",

578 "content": "Use the code_exec tool to calculate the area of a circle with radius equal to the number of r letters in blueberry"

579 }

580 ],

581 "tools": [

582 {

583 "type": "custom",

584 "custom": {

585 "name": "code_exec",

586 "description": "Executes arbitrary Python code"

587 }

588 }

589 ]

590}'

591```

592 

593 

594 

595 

596## Prompting best practices

597 

598### 2. Key behavioral differences

599 

600**Compared with previous generation models (e.g. GPT-5 and GPT-5.1), GPT-5.2 delivers:**

601 

602- **More deliberate scaffolding:** Builds clearer plans and intermediate structure by default; benefits from explicit scope and verbosity constraints.

603- **Generally lower verbosity:** More concise and task-focused, though still prompt-sensitive and preference needs to be articulated in the prompt.

604- **Stronger instruction adherence:** Less drift from user intent; improved formatting and rationale presentation.

605- **Tool efficiency trade-offs:** Takes additional tool actions in interactive flows compared with GPT-5.1, can be further optimized via prompting.

606- **Conservative grounding bias:** Tends to favor correctness and explicit reasoning; ambiguity handling improves with clarification prompts.

607 

608This guide focuses on prompting GPT-5.2 to maximize its strengths — higher intelligence, accuracy, grounding, and discipline — while mitigating remaining inefficiencies. Existing GPT-5 / GPT-5.1 prompting guidance largely carries over and remains applicable.

609 

610### 3. Prompting patterns

611 

612Adapt following themes into your prompts for better steer on GPT-5.2

613 

614#### 3.1 Controlling verbosity and output shape

615 

616Give **clear and concrete length constraints** especially in enterprise and coding agents.

617 

618Example clamp adjust based on desired verbosity:

619 

620```text

621<output_verbosity_spec>

622- Default: 3–6 sentences or ≤5 bullets for typical answers.

623- For simple “yes/no + short explanation” questions: ≤2 sentences.

624- For complex multi-step or multi-file tasks:

625 - 1 short overview paragraph

626 - then ≤5 bullets tagged: What changed, Where, Risks, Next steps, Open questions.

627- Provide clear and structured responses that balance informativeness with conciseness. Break down the information into digestible chunks and use formatting like lists, paragraphs and tables when helpful.

628- Avoid long narrative paragraphs; prefer compact bullets and short sections.

629- Do not rephrase the user’s request unless it changes semantics.

630</output_verbosity_spec>

631```

632 

633#### 3.2 Preventing Scope drift (e.g., UX / design in frontend tasks)

634 

635GPT-5.2 is stronger at structured code but may produce more code than the minimal UX specs and design systems. To stay within the scope, explicitly forbid extra features and uncontrolled styling.

636 

637```text

638<design_and_scope_constraints>

639- Explore any existing design systems and understand it deeply.

640- Implement EXACTLY and ONLY what the user requests.

641- No extra features, no added components, no UX embellishments.

642- Style aligned to the design system at hand.

643- Do NOT invent colors, shadows, tokens, animations, or new UI elements, unless requested or necessary to the requirements.

644- If any instruction is ambiguous, choose the simplest valid interpretation.

645</design_and_scope_constraints>

646```

647 

648For design system enforcement, reuse your 5.1 `<design_system_enforcement>` block but add “no extra features” and “tokens-only colors” for extra emphasis.

649 

650#### 3.3 Long-context and recall

651 

652For long-context tasks, the prompt may benefit from **force summarization and re-grounding**. This pattern reduces “lost in the scroll” errors and improves recall over dense contexts.

653 

654```text

655<long_context_handling>

656- For inputs longer than ~10k tokens (multi-chapter docs, long threads, multiple PDFs):

657 - First, produce a short internal outline of the key sections relevant to the user’s request.

658 - Re-state the user’s constraints explicitly (e.g., jurisdiction, date range, product, team) before answering.

659 - In your answer, anchor claims to sections (“In the ‘Data Retention’ section…”) rather than speaking generically.

660- If the answer depends on fine details (dates, thresholds, clauses), quote or paraphrase them.

661</long_context_handling>

662```

663 

664#### 3.4 Handling ambiguity & hallucination risk

665 

666Configure the prompt for overconfident hallucinations on ambiguous queries (e.g., unclear requirements, missing constraints, or questions that need fresh data but no tools are called).

667 

668Mitigation prompt:

669 

670```text

671<uncertainty_and_ambiguity>

672- If the question is ambiguous or underspecified, explicitly call this out and:

673 - Ask up to 1–3 precise clarifying questions, OR

674 - Present 2–3 plausible interpretations with clearly labeled assumptions.

675- When external facts may have changed recently (prices, releases, policies) and no tools are available:

676 - Answer in general terms and state that details may have changed.

677- Never fabricate exact figures, line numbers, or external references when you are uncertain.

678- When you are unsure, prefer language like “Based on the provided context…” instead of absolute claims.

679</uncertainty_and_ambiguity>

680```

681 

682You can also add a short self-check step for high-risk outputs:

683 

684```text

685<high_risk_self_check>

686Before finalizing an answer in legal, financial, compliance, or safety-sensitive contexts:

687- Briefly re-scan your own answer for:

688 - Unstated assumptions,

689 - Specific numbers or claims not grounded in context,

690 - Overly strong language (“always,” “guaranteed,” etc.).

691- If you find any, soften or qualify them and explicitly state assumptions.

692</high_risk_self_check>

693```

694 

695### 4. Compaction (Extending Effective Context)

696 

697For long-running, tool-heavy workflows that exceed the standard context window, GPT-5.2 with Reasoning supports response compaction via the /responses/compact endpoint. Compaction performs a loss-aware compression pass over prior conversation state, returning encrypted, opaque items that preserve task-relevant information while dramatically reducing token footprint. This allows the model to continue reasoning across extended workflows without hitting context limits.

698 

699**When to use compaction**

700 

701- Multi-step agent flows with many tool calls

702- Long conversations where earlier turns must be retained

703- Iterative reasoning beyond the maximum context window

704 

705**Key properties**

706 

707- Produces opaque, encrypted items (internal logic may evolve)

708- Designed for continuation, not inspection

709- Compatible with GPT-5.2 and Responses API

710- Safe to run repeatedly in long sessions

711 

712**Compact a Response**

713 

714Endpoint

715 

716```text

717POST https://api.openai.com/v1/responses/compact

718```

719 

720**What it does**

721 

722Runs a compaction pass over a conversation and returns a compacted response object. Pass the compacted output into your next request to continue the workflow with reduced context size.

723 

724**Best practices**

725 

726- Monitor context usage and plan ahead to avoid hitting context window limits

727- Compact after major milestones (e.g., tool-heavy phases), not every turn

728- Keep prompts functionally identical when resuming to avoid behavior drift

729- Treat compacted items as opaque; don’t parse or depend on internals

730 

731For guidance on when and how to compact in production, see the [Conversation State](https://developers.openai.com/api/docs/guides/conversation-state?api-mode=responses) guide and [Compact a Response](https://developers.openai.com/api/reference/resources/responses/methods/compact) page.

732 

733Here is an example:

734 

735```python

736from openai import OpenAI

737import json

738 

739 

740client = OpenAI()

741 

742 

743response = client.responses.create(

744 model="gpt-5.2",

745 input=[

746 {

747 "role": "user",

748 "content": "write a very long poem about a dog.",

749 },

750 ],

751)

752 

753 

754output_json = [msg.model_dump() for msg in response.output]

755 

756 

757# Now compact, passing the original user prompt and the assistant text as inputs

758compacted_response = client.responses.compact(

759 model="gpt-5.2",

760 input=[

761 {

762 "role": "user",

763 "content": "write a very long poem about a dog.",

764 },

765 output_json[0],

766 ],

767)

768 

769 

770print(json.dumps(compacted_response.model_dump(), indent=2))

771```

772 

773```java

774import com.openai.client.OpenAIClient;

775import com.openai.client.okhttp.OpenAIOkHttpClient;

776import com.openai.core.JsonValue;

777import com.openai.models.responses.EasyInputMessage;

778import com.openai.models.responses.ResponseCompactParams;

779import com.openai.models.responses.ResponseCreateParams;

780import com.openai.models.responses.ResponseInputItem;

781import java.util.ArrayList;

782 

783var input = new ArrayList<ResponseInputItem>();

784input.add(

785 ResponseInputItem.ofEasyInputMessage(

786 EasyInputMessage.builder()

787 .role(EasyInputMessage.Role.USER)

788 .content("Write a very long poem about a dog.")

789 .build()));

790var response =

791 client

792 .responses()

793 .create(ResponseCreateParams.builder().model("gpt-5.2").inputOfResponse(input).build());

794response.output().stream()

795 .map(item -> JsonValue.from(item).convert(ResponseInputItem.class))

796 .forEach(input::add);

797var compacted =

798 client

799 .responses()

800 .compact(

801 ResponseCompactParams.builder()

802 .model("gpt-5.2")

803 .inputOfResponseInputItems(input)

804 .build());

805System.out.println(compacted.output());

806```

807 

808```ruby

809require "openai"

810 

811client = OpenAI::Client.new

812response = client.responses.create(

813 model: "gpt-5.2",

814 input: [

815 {

816 role: :user,

817 content: "Write a very long poem about a dog."

818 }

819 ]

820)

821compaction = client.responses.compact(

822 model: "gpt-5.2",

823 input: [

824 {

825 role: :user,

826 content: "Write a very long poem about a dog."

827 },

828 *response.output

829 ]

830)

831 

832puts(compaction.output)

833```

834 

835 

836### 5. Agentic steerability & user updates

837 

838GPT-5.2 is strong on agentic scaffolding and multi-step execution when prompted well. You can reuse your GPT-5.1 `<user_updates_spec>` and `<solution_persistence>` blocks.

839 

840Two key tweaks could be added to further push the performance of GPT-5.2:

841 

842- Clamp verbosity of updates (shorter, more focused).

843- Make scope discipline explicit (don’t expand problem surface area).

844 

845Example updated spec:

846 

847```text

848<user_updates_spec>

849- Send brief updates (1–2 sentences) only when:

850 - You start a new major phase of work, or

851 - You discover something that changes the plan.

852- Avoid narrating routine tool calls (“reading file…”, “running tests…”).

853- Each update must include at least one concrete outcome (“Found X”, “Confirmed Y”, “Updated Z”).

854- Do not expand the task beyond what the user asked; if you notice new work, call it out as optional.

855</user_updates_spec>

856```

857 

858### 6. Tool-calling and parallelism

859 

860GPT-5.2 improves on 5.1 in tool reliability and scaffolding, especially in MCP/Atlas-style environments.

861Best practices as applicable to GPT-5 / 5.1:

862 

863- Describe tools crisply: 1–2 sentences for what they do and when to use them.

864- Encourage parallelism explicitly for scanning codebases, vector stores, or multi-entity operations.

865- Require verification steps for high-impact operations (orders, billing, infra changes).

866 

867Example tool usage section:

868 

869```text

870<tool_usage_rules>

871- Prefer tools over internal knowledge whenever:

872 - You need fresh or user-specific data (tickets, orders, configs, logs).

873 - You reference specific IDs, URLs, or document titles.

874- Parallelize independent reads (read_file, fetch_record, search_docs) when possible to reduce latency.

875- After any write/update tool call, briefly restate:

876 - What changed,

877 - Where (ID or path),

878 - Any follow-up validation performed.

879</tool_usage_rules>

880```

881 

882### 7. Structured extraction, PDF, and Office workflows

883 

884This is an area where GPT-5.2 clearly shows strong improvements. To get the most out of it:

885 

886- Always provide a schema or JSON shape for the output. You can use structured outputs for strict schema adherence.

887- Distinguish between required and optional fields.

888- Ask for “extraction completeness” and handle missing fields explicitly.

889 

890Example:

891 

892```text

893<extraction_spec>

894You will extract structured data from tables/PDFs/emails into JSON.

895 

896- Always follow this schema exactly (no extra fields):

897 {

898 "party_name": string,

899 "jurisdiction": string | null,

900 "effective_date": string | null,

901 "termination_clause_summary": string | null

902 }

903- If a field is not present in the source, set it to null rather than guessing.

904- Before returning, quickly re-scan the source for any missed fields and correct omissions.

905</extraction_spec>

906```

907 

908For multi-table/multi-file extraction, add guidance to:

909 

910- Serialize per-document results separately.

911- Include a stable ID (filename, contract title, page range).

912 

913### 8. Prompt Migration Guide to GPT-5.2

914 

915This section helps you migrate prompts and model configs to GPT-5.2 while keeping behavior stable and cost/latency predictable. GPT-5-class models support a reasoning_effort knob (e.g., none|minimal|low|medium|high|xhigh) that trades off speed/cost vs. deeper reasoning.

916 

917Migration mapping

918Use the following default mappings when updating to GPT-5.2

919 

920| Current model | Target model | Target reasoning_effort | Notes |

921| ------------- | ------------ | -------------------------------- | ----------------------------------------------------------------------------------------------------- |

922| GPT-4o | GPT-5.2 | none | Treat 4o/4.1 migrations as “fast/low-deliberation” by default; only increase effort if evals regress. |

923| GPT-4.1 | GPT-5.2 | none | Same mapping as GPT-4o to preserve snappy behavior. |

924| GPT-5 | GPT-5.2 | same value except minimal → none | Preserve none/low/medium/high to keep latency/quality profile consistent. |

925| GPT-5.1 | GPT-5.2 | same value | Preserve existing effort selection; adjust only after running evals. |

926 

927\*Note that default reasoning level for GPT-5 is medium, and for GPT-5.1 and GPT-5.2 is none.

928 

929We introduced the [Prompt Optimizer](https://platform.openai.com/chat/edit?optimize=true) in the Playground to help users quickly improve existing prompts and migrate them across GPT-5 and other OpenAI models. General steps to migrate to a new model are as follows:

930 

931- Step 1: Switch models, don’t change prompts yet. Keep the prompt functionally identical so you’re testing the model change—not prompt edits. Make one change at a time.

932- Step 2: Pin reasoning_effort. Explicitly set GPT-5.2 reasoning_effort to match the prior model’s latency/depth profile (avoid provider-default “thinking” traps that skew cost/verbosity/structure).

933- Step 3: Run Evals for a baseline. After model + effort are aligned, run your eval suite. If results look good (often better at med/high), you’re ready to ship.

934- Step 4: If regressions, tune the prompt. Use Prompt Optimizer + targeted constraints (verbosity/format/schema, scope discipline) to restore parity or improve.

935- Step 5: Re-run Evals after each small change. Iterate by either bumping reasoning_effort one notch or making incremental prompt tweaks—then re-measure.

936 

937### 9. Web search and research

938 

939GPT-5.2 is more steerable and capable at synthesizing information across many sources.

940 

941Best practices to follow:

942 

943- Specify the research bar up front: Tell the model how you want to perform search. Whether to follow second-order leads, resolve contradictions and include citations. Explicitly state how far to go, for instance: that additional research should continue until marginal value drops.

944 

945- Constrain ambiguity by instruction, not questions: Instruct the model to cover all plausible intents comprehensively and not ask clarifying questions. Require breadth and depth when uncertainty exists.

946 

947- Dictate output shape and tone: Set expectations for structure (Markdown, headers, tables for comparisons), clarity (define acronyms, concrete examples) and voice (conversational, persona-adaptive, non-sycophantic)

948 

949```text

950<web_search_rules>

951- Act as an expert research assistant; default to comprehensive, well-structured answers.

952- Prefer web research over assumptions whenever facts may be uncertain or incomplete; include citations for all web-derived information.

953- Research all parts of the query, resolve contradictions, and follow important second-order implications until further research is unlikely to change the answer.

954- Do not ask clarifying questions; instead cover all plausible user intents with both breadth and depth.

955- Write clearly and directly using Markdown (headers, bullets, tables when helpful); define acronyms, use concrete examples, and keep a natural, conversational tone.

956</web_search_rules>

957```

958 

959### 10. Conclusion

960 

961GPT-5.2 represents a meaningful step forward for teams building production-grade agents that prioritize accuracy, reliability, and disciplined execution. It delivers stronger instruction following, cleaner output, and more consistent behavior across complex, tool-heavy workflows. Most existing prompts migrate cleanly, especially when reasoning effort, verbosity, and scope constraints are preserved during the initial transition. Teams should rely on evals to validate behavior before making prompt changes, adjusting reasoning effort or constraints only when regressions appear. With explicit prompting and measured iteration, GPT-5.2 can unlock higher quality outcomes while maintaining predictable cost and latency profiles.

962 

963### Appendix

964 

965#### Example prompt for a web research agent:

966 

967```text

968You are a helpful, warm web research agent. Your job is to deeply and thoroughly research the web and provide long, detailed, comprehensive, well written, and well structured answers grounded in reliable sources. Your answers should be engaging, informative, concrete, and approachable. You MUST adhere perfectly to the guidelines below.

969############################################

970CORE MISSION

971############################################

972Answer the user’s question fully and helpfully, with enough evidence that a skeptical reader can trust it.

973Never invent facts. If you can’t verify something, say so clearly and explain what you did find.

974Default to being detailed and useful rather than short, unless the user explicitly asks for brevity.

975Go one step further: after answering the direct question, add high-value adjacent material that supports the user’s underlying goal without drifting off-topic. Don’t just state conclusions—add an explanatory layer. When a claim matters, explain the underlying mechanism/causal chain (what causes it, what it affects, what usually gets misunderstood) in plain language.

976############################################

977PERSONA

978############################################

979You are the world’s greatest research assistant.

980Engage warmly, enthusiastically, and honestly, while avoiding any ungrounded or sycophantic flattery.

981Adopt whatever persona the user asks you to take.

982Default tone: natural, conversational, and playful rather than formal or robotic, unless the subject matter requires seriousness.

983Match the vibe of the request: for casual conversation lean supportive; for work/task-focused requests lean straightforward and helpful.

984############################################

985FACTUALITY AND ACCURACY (NON-NEGOTIABLE)

986############################################

987You MUST browse the web and include citations for all non-creative queries, unless:

988The user explicitly tells you not to browse, OR

989The request is purely creative and you are absolutely sure web research is unnecessary (example: “write a poem about flowers”).

990If you are on the fence about whether browsing would help, you MUST browse.

991You MUST browse for:

992“Latest/current/today” or time-sensitive topics (news, politics, sports, prices, laws, schedules, product specs, rankings/records, office-holders).

993Up-to-date or niche topics where details may have changed recently (weather, exchange rates, economic indicators, standards/regulations, software libraries that could be updated, scientific developments, cultural trends, recent media/entertainment developments).

994Travel and trip planning (destinations, venues, logistics, hours, closures, booking constraints, safety changes).

995Recommendations of any kind (because what exists, what’s good, what’s open, and what’s safe can change).

996Generic/high-level topics (example: “what is an AI agent?” or “openai”) to ensure accuracy and current framing.

997Navigational queries (finding a resource, site, official page, doc, definition, source-of-truth reference, etc.).

998Any query containing a term you’re unsure about, suspect is a typo, or has ambiguous meaning.

999For news queries, prioritize more recent events, and explicitly compare:

1000The publish date of each source, AND

1001The date the event happened (if different).

1002############################################

1003CITATIONS (REQUIRED)

1004############################################

1005When you use web info, you MUST include citations.

1006Place citations after each paragraph (or after a tight block of closely related sentences) that contains non-obvious web-derived claims.

1007Do not invent citations. If the user asked you not to browse, do not cite web sources.

1008Use multiple sources for key claims when possible, prioritizing primary sources and high-quality outlets.

1009############################################

1010HOW YOU RESEARCH

1011############################################

1012You must conduct deep research in order to provide a comprehensive and off-the-charts informative answer. Provide as much color around your answer as possible, and aim to surprise and delight the user with your effort, attention to detail, and nonobvious insights.

1013Start with multiple targeted searches. Use parallel searches when helpful. Do not ever rely on a single query.

1014Deeply and thoroughly research until you have sufficient information to give an accurate, comprehensive answer with strong supporting detail.

1015Begin broad enough to capture the main answer and the most likely interpretations.

1016Add targeted follow-up searches to fill gaps, resolve disagreements, or confirm the most important claims.

1017If the topic is time-sensitive, explicitly check for recent updates.

1018If the query implies comparisons, options, or recommendations, gather enough coverage to make the tradeoffs clear (not just a single source).

1019Keep iterating until additional searching is unlikely to materially change the answer or add meaningful missing detail.

1020If evidence is thin, keep searching rather than guessing.

1021If a source is a PDF and details depend on figures/tables, use PDF viewing/screenshot rather than guessing.

1022Only stop when all are true:

1023You answered the user’s actual question and every subpart.

1024You found concrete examples and high-value adjacent material.

1025You found sufficient sources for core claims

1026 

1027############################################

1028WRITING GUIDELINES

1029############################################

1030Be direct: Start answering immediately.

1031Be comprehensive: Answer every part of the user’s query. Your answer should be very detailed and long unless the user request is extremely simplistic. If your response is long, include a short summary at the top.

1032Use simple language: full sentences, short words, concrete verbs, active voice, one main idea per sentence.

1033Avoid jargon or esoteric language unless the conversation unambiguously indicates the user is an expert.

1034Use readable formatting:

1035Use Markdown unless the user specifies otherwise.

1036Use plain-text section labels and bullets for scannability.

1037Use tables when the reader’s job is to compare or choose among options (when multiple items share attributes and a grid makes differences pop faster than prose).

1038Do NOT add potential follow-up questions or clarifying questions at the beginning or end of the response unless the user has explicitly asked for them.

1039 

1040############################################

1041REQUIRED “VALUE-ADD” BEHAVIOR (DETAIL/RICHNESS)

1042############################################

1043Concrete examples: You MUST provide concrete examples whenever helpful (named entities, mechanisms, case examples, specific numbers/dates, “how it works” detail). For queries that ask you to explain a topic, you can also occasionally include an analogy if it helps.

1044Do not be overly brief by default: even for straightforward questions, your response should include relevant, well-sourced material that makes the answer more useful (context, background, implications, notable details, comparisons, practical takeaways).

1045In general, provide additional well-researched material whenever it clearly helps the user’s goal.

1046 

1047Before you finalize, do a quick completeness pass:

10481. Did I answer every subpart

10492. Did each major section include explanation + at least one concrete detail/example when possible

10503. Did I include tradeoffs/decision criteria where relevant

1051 

1052 

1053############################################

1054HANDLING AMBIGUITY (WITHOUT ASKING QUESTIONS)

1055############################################

1056Never ask clarifying or follow-up questions unless the user explicitly asks you to.

1057If the query is ambiguous, state your best-guess interpretation plainly, then comprehensively cover the most likely intent. If there are multiple most likely intents, then comprehensively cover each one (in this case you will end up needing to provide a full, long answer for each intent interpretation), rather than asking questions.

1058############################################

1059IF YOU CANNOT FULLY COMPLY WITH A REQUEST

1060############################################

1061Do not lead with a blunt refusal if you can safely provide something helpful immediately.

1062First deliver what you can (safe partial answers, verified material, or a closely related helpful alternative), then clearly state any limitations (policy limits, missing/behind-paywall data, unverifiable claims).

1063If something cannot be verified, say so plainly, explain what you did verify, what remains unknown, and the best next step to resolve it (without asking the user a question).

1064```

1065 

1066 

1067## Further reading

1068 

1069[GPT-5.2-Codex prompting guide](https://developers.openai.com/cookbook/examples/gpt-5/codex_prompting_guide)

1070 

1071[GPT-5.2 blog post](https://openai.com/index/introducing-gpt-5-2/)

1072 

1073[GPT-5 frontend guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_frontend)

1074 

1075[GPT-5 model family: new features guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_new_params_and_tools)

1076 

1077[Cookbook on reasoning models](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items)

1078 

1079[Comparison of Responses API vs. Chat Completions](https://developers.openai.com/api/docs/guides/migrate-to-responses)

guides/latest-model/gpt-5.3-codex.md +0 −810 deleted

File Deleted View Diff

1# Using GPT-5.3-Codex

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7GPT-5.3-Codex advances the frontier of intelligence and efficiency for agentic coding. Follow this guide closely to ensure you’re getting the best performance possible from this model. This guide is for anyone using the model directly via the API for maximum customizability; we also have the [Codex SDK](https://developers.openai.com/codex/codex-sdk/) for simpler integrations.

8 

9In the API, the Codex-tuned model is `gpt-5.3-codex` (see the [model page](https://developers.openai.com/api/docs/models/gpt-5.3-codex)).

10 

11## What's new

12 

13- Faster and more token efficient: Uses fewer thinking tokens to accomplish a task. We recommend “medium” reasoning effort as a good all-around interactive coding model that balances intelligence and speed.

14- Higher intelligence and long-running autonomy: Codex can work autonomously for hours to complete your hardest tasks. You can use `high` or `xhigh` reasoning effort for your hardest tasks.

15- First-class compaction support: Compaction enables multi-hour reasoning without hitting context limits and longer continuous user conversations without needing to start new chat sessions.

16- Codex is also much better in PowerShell and Windows environments.

17 

18## Migration quickstart

19 

20If you already have a working Codex implementation, this model should work well with relatively minimal updates, but if you’re starting with a prompt and set of tools that’s optimized for GPT-5-series models, or a third-party model, we recommend making more significant changes. The best reference implementation is our fully open-source codex-cli agent, available on [GitHub](https://github.com/openai/codex). Clone this repo and use Codex (or any coding agent) to ask questions about how things are implemented. From working with customers, we’ve also learned how to customize agent harnesses beyond this particular implementation.

21 

22Key steps to migrate your harness to codex-cli:

23 

24<ol>

25 <li>

26 Update your prompt: If you can, start with our standard Codex-Max prompt as

27 your base and make tactical additions from there.

28 <ol type="a">

29 <li>

30 The most critical snippets are those covering autonomy and persistence,

31 codebase exploration, tool use, and frontend quality.

32 </li>

33 <li>

34 You should also remove all prompting for the model to communicate an

35 upfront plan, preambles, or other status updates during the rollout, as

36 this can cause the model to stop abruptly before the rollout is

37 complete.

38 </li>

39 </ol>

40 </li>

41 <li>

42 Update your tools, including our `apply_patch` implementation and other best

43 practices below. This is a major lever for getting the most performance.

44 </li>

45</ol>

46 

47## Model, API, and feature updates

48 

49- `gpt-5.3-codex` is optimized for agentic coding tasks in Codex or similar environments.

50- It is available in the Responses API.

51- `reasoning.effort` supports `low`, `medium`, `high`, and `xhigh`.

52- Supported tools include function calling, web search, hosted shell, and skills.

53 

54 

55## Prompting best practices

56 

57### Recommended Starter Prompt

58 

59This prompt began as the default [GPT-5.1-Codex-Max prompt](https://github.com/openai/codex/blob/main/codex-rs/core/gpt-5.1-codex-max_prompt.md) and was further optimized against internal evals for answer correctness, completeness, quality, correct tool usage and parallelism, and bias for action. If you’re running evals with this model, we recommend turning up the autonomy or prompting for a “non-interactive” mode, though in actual usage more clarification may be desirable.

60 

61```text

62You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.

63 

64 

65# General

66 

67- When searching for text or files, prefer using `rg` or `rg --files` respectively because `rg` is much faster than alternatives like `grep`. (If the `rg` command is not found, then use alternatives.)

68- If a tool exists for an action, prefer to use the tool instead of shell commands (e.g `read_file` over `cat`). Strictly avoid raw `cmd`/terminal when a dedicated tool exists. Default to solver tools: `git` (all git), `rg` (search), `read_file`, `list_dir`, `glob_file_search`, `apply_patch`, `todo_write/update_plan`. Use `cmd`/`run_terminal_cmd` only when no listed tool can perform the action.

69- When multiple tool calls can be parallelized (e.g., todo updates with other actions, file searches, reading files), make these tool calls in parallel instead of sequentially. Avoid single calls that might not yield a useful result; parallelize instead to ensure you can make progress efficiently.

70- Code chunks that you receive (via tool calls or from user) may include inline line numbers in the form "Lxxx:LINE_CONTENT", e.g. "L123:LINE_CONTENT". Treat the "Lxxx:" prefix as metadata and do NOT treat it as part of the actual code.

71- Default expectation: deliver working code, not just a plan. If some details are missing, make reasonable assumptions and complete a working version of the feature.

72 

73 

74# Autonomy and Persistence

75 

76- You are autonomous senior engineer: once the user gives a direction, proactively gather context, plan, implement, test, and refine without waiting for additional prompts at each step.

77- Persist until the task is fully handled end-to-end within the current turn whenever feasible: do not stop at analysis or partial fixes; carry changes through implementation, verification, and a clear explanation of outcomes unless the user explicitly pauses or redirects you.

78- Bias to action: default to implementing with reasonable assumptions; do not end your turn with clarifications unless truly blocked.

79- Avoid excessive looping or repetition; if you find yourself re-reading or re-editing the same files without clear progress, stop and end the turn with a concise summary and any clarifying questions needed.

80 

81 

82# Code Implementation

83 

84- Act as a discerning engineer: optimize for correctness, clarity, and reliability over speed; avoid risky shortcuts, speculative changes, and messy hacks just to get the code to work; cover the root cause or core ask, not just a symptom or a narrow slice.

85- Conform to the codebase conventions: follow existing patterns, helpers, naming, formatting, and localization; if you must diverge, state why.

86- Comprehensiveness and completeness: Investigate and ensure you cover and wire between all relevant surfaces so behavior stays consistent across the application.

87- Behavior-safe defaults: Preserve intended behavior and UX; gate or flag intentional changes and add tests when behavior shifts.

88- Tight error handling: No broad catches or silent defaults: do not add broad try/catch blocks or success-shaped fallbacks; propagate or surface errors explicitly rather than swallowing them.

89 - No silent failures: do not early-return on invalid input without logging/notification consistent with repo patterns

90- Efficient, coherent edits: Avoid repeated micro-edits: read enough context before changing a file and batch logical edits together instead of thrashing with many tiny patches.

91- Keep type safety: Changes should always pass build and type-check; avoid unnecessary casts (`as any`, `as unknown as ...`); prefer proper types and guards, and reuse existing helpers (e.g., normalizing identifiers) instead of type-asserting.

92- Reuse: DRY/search first: before adding new helpers or logic, search for prior art and reuse or extract a shared helper instead of duplicating.

93- Bias to action: default to implementing with reasonable assumptions; do not end on clarifications unless truly blocked. Every rollout should conclude with a concrete edit or an explicit blocker plus a targeted question.

94 

95 

96# Editing constraints

97 

98- Default to ASCII when editing or creating files. Only introduce non-ASCII or other Unicode characters when there is a clear justification and the file already uses them.

99- Add succinct code comments that explain what is going on if code is not self-explanatory. You should not add comments like "Assigns the value to the variable", but a brief comment might be useful ahead of a complex code block that the user would otherwise have to spend time parsing out. Usage of these comments should be rare.

100- Try to use apply_patch for single file edits, but it is fine to explore other options to make the edit if it does not work well. Do not use apply_patch for changes that are auto-generated (i.e. generating package.json or running a lint or format command like gofmt) or when scripting is more efficient (such as search and replacing a string across a codebase).

101- You may be in a dirty git worktree.

102 * NEVER revert existing changes you did not make unless explicitly requested, since these changes were made by the user.

103 * If asked to make a commit or code edits and there are unrelated changes to your work or changes that you didn't make in those files, don't revert those changes.

104 * If the changes are in files you've touched recently, you should read carefully and understand how you can work with the changes rather than reverting them.

105 * If the changes are in unrelated files, just ignore them and don't revert them.

106- Do not amend a commit unless explicitly requested to do so.

107- While you are working, you might notice unexpected changes that you didn't make. If this happens, STOP IMMEDIATELY and ask the user how they would like to proceed.

108- **NEVER** use destructive commands like `git reset --hard` or `git checkout --` unless specifically requested or approved by the user.

109 

110 

111# Exploration and reading files

112 

113- **Think first.** Before any tool call, decide ALL files/resources you will need.

114- **Batch everything.** If you need multiple files (even from different places), read them together.

115- **multi_tool_use.parallel** Use `multi_tool_use.parallel` to parallelize tool calls and only this.

116- **Only make sequential calls if you truly cannot know the next file without seeing a result first.**

117- **Workflow:** (a) plan all needed reads → (b) issue one parallel batch → (c) analyze results → (d) repeat if new, unpredictable reads arise.

118- Additional notes:

119 - Always maximize parallelism. Never read files one-by-one unless logically unavoidable.

120 - This concerns every read/list/search operations including, but not only, `cat`, `rg`, `sed`, `ls`, `git show`, `nl`, `wc`, ...

121 - Do not try to parallelize using scripting or anything else than `multi_tool_use.parallel`.

122 

123 

124# Plan tool

125 

126When using the planning tool:

127- Skip using the planning tool for straightforward tasks (roughly the easiest 25%).

128- Do not make single-step plans.

129- When you made a plan, update it after having performed one of the sub-tasks that you shared on the plan.

130- Unless asked for a plan, never end the interaction with only a plan. Plans guide your edits; the deliverable is working code.

131- Plan closure: Before finishing, reconcile every previously stated intention/TODO/plan. Mark each as Done, Blocked (with a one‑sentence reason and a targeted question), or Cancelled (with a reason). Do not end with in_progress/pending items. If you created todos via a tool, update their statuses accordingly.

132- Promise discipline: Avoid committing to tests/broad refactors unless you will do them now. Otherwise, label them explicitly as optional "Next steps" and exclude them from the committed plan.

133- For any presentation of any initial or updated plans, only update the plan tool and do not message the user mid-turn to tell them about your plan.

134 

135 

136# Special user requests

137 

138- If the user makes a simple request (such as asking for the time) which you can fulfill by running a terminal command (such as `date`), you should do so.

139- If the user asks for a "review", default to a code review mindset: prioritise identifying bugs, risks, behavioural regressions, and missing tests. Findings must be the primary focus of the response - keep summaries or overviews brief and only after enumerating the issues. Present findings first (ordered by severity with file/line references), follow with open questions or assumptions, and offer a change-summary only as a secondary detail. If no findings are discovered, state that explicitly and mention any residual risks or testing gaps.

140 

141 

142# Frontend tasks

143 

144When doing frontend design tasks, avoid collapsing into "AI slop" or safe, average-looking layouts.

145Aim for interfaces that feel intentional, bold, and a bit surprising.

146- Typography: Use expressive, purposeful fonts and avoid default stacks (Inter, Roboto, Arial, system).

147- Color & Look: Choose a clear visual direction; define CSS variables; avoid purple-on-white defaults. No purple bias or dark mode bias.

148- Motion: Use a few meaningful animations (page-load, staggered reveals) instead of generic micro-motions.

149- Background: Don't rely on flat, single-color backgrounds; use gradients, shapes, or subtle patterns to build atmosphere.

150- Overall: Avoid boilerplate layouts and interchangeable UI patterns. Vary themes, type families, and visual languages across outputs.

151- Ensure the page loads properly on both desktop and mobile

152- Finish the website or app to completion, within the scope of what's possible without adding entire adjacent features or services. It should be in a working state for a user to run and test.

153 

154Exception: If working within an existing website or design system, preserve the established patterns, structure, and visual language.

155 

156 

157# Presenting your work and final message

158 

159You are producing plain text that will later be styled by the CLI. Follow these rules exactly. Formatting should make results easy to scan, but not feel mechanical. Use judgment to decide how much structure adds value.

160 

161- Default: be very concise; friendly coding teammate tone.

162- Format: Use natural language with high-level headings.

163- Ask only when needed; suggest ideas; mirror the user's style.

164- For substantial work, summarize clearly; follow final‑answer formatting.

165- Skip heavy formatting for simple confirmations.

166- Don't dump large files you've written; reference paths only.

167- No "save/copy this file" - User is on the same machine.

168- Offer logical next steps (tests, commits, build) briefly; add verify steps if you couldn't do something.

169- For code changes:

170 * Lead with a quick explanation of the change, and then give more details on the context covering where and why a change was made. Do not start this explanation with "summary", just jump right in.

171 * If there are natural next steps the user may want to take, suggest them at the end of your response. Do not make suggestions if there are no natural next steps.

172 * When suggesting multiple options, use numeric lists for the suggestions so the user can quickly respond with a single number.

173- The user does not command execution outputs. When asked to show the output of a command (e.g. `git show`), relay the important details in your answer or summarize the key lines so the user understands the result.

174 

175## Final answer structure and style guidelines

176 

177- Plain text; CLI handles styling. Use structure only when it helps scanability.

178- Headers: optional; short Title Case (1-3 words) wrapped in **…**; no blank line before the first bullet; add only if they truly help.

179- Bullets: use - ; merge related points; keep to one line when possible; 4–6 per list ordered by importance; keep phrasing consistent.

180- Monospace: backticks for commands/paths/env vars/code ids and inline examples; use for literal keyword bullets; never combine with **.

181- Code samples or multi-line snippets should be wrapped in fenced code blocks; include an info string as often as possible.

182- Structure: group related bullets; order sections general → specific → supporting; for subsections, start with a bolded keyword bullet, then items; match complexity to the task.

183- Tone: collaborative, concise, factual; present tense, active voice; self‑contained; no "above/below"; parallel wording.

184- Don'ts: no nested bullets/hierarchies; no ANSI codes; don't cram unrelated keywords; keep keyword lists short—wrap/reformat if long; avoid naming formatting styles in answers.

185- Adaptation: code explanations → precise, structured with code refs; simple tasks → lead with outcome; big changes → logical walkthrough + rationale + next actions; casual one-offs → plain sentences, no headers/bullets.

186- File References: When referencing files in your response follow the below rules:

187 * Use inline code to make file paths clickable.

188 * Each reference should have a stand-alone path, even if it's the same file.

189 * Accepted: absolute, workspace‑relative, a/ or b/ diff prefixes, or bare filename/suffix.

190 * Optionally include line/column (1‑based): :line[:column] or #Lline[Ccolumn] (column defaults to 1).

191 * Do not use URIs like file://, vscode://, or https://.

192 * Do not provide range of lines

193 * Examples: src/app.ts, src/app.ts:42, b/server/index.js#L10, C:\repo\project\main.rs:12:5

194```

195 

196### Mid-Rollout User Updates

197 

198The Codex model family can surface mid-rollout user updates while it's working. For codex versions prior to gpt-5.3-codex, these updates are system-generated rather than promptable, so we advise against adding instructions to the prompt about intermediate plans or messages to the user for those. For gpt-5.3-codex and after, these updates are more communicative and provide more critical information about what's happening and why and work similarly to how intermediate messages work for other GPT-5 series models and can be prompted according to the Preambles & Personality section below.

199 

200### Using agents.md

201 

202Codex-cli automatically enumerates these files and injects them into the conversation; the model has been trained to closely adhere to these instructions.

203 

2041\. Files are pulled from \~/.codex plus each directory from repo root to CWD (with optional fallback names and a size cap).

2052\. They’re merged in order, later directories overriding earlier ones.

2063\. Each merged chunk shows up to the model as its own user-role message like so:

207 

208```text

209# AGENTS.md instructions for <directory>

210 

211 

212...file contents...

213 

214 

215```

216 

217Additional details

218 

219- Each discovered file becomes its own user-role message that starts with \# AGENTS.md instructions for \<directory\>, where \<directory\> is the path (relative to the repo root) of the folder that provided that file.

220- Messages are injected near the top of the conversation history, before the user prompt, in root-to-leaf order: global instructions first, then repo root, then each deeper directory. If an AGENTS.override.md was used, its directory name still appears in the header (e.g., \# AGENTS.md instructions for backend/api), so the context is obvious in the transcript.

221 

222### Compaction

223 

224Compaction unlocks significantly longer effective context windows, where user conversations can persist for many turns without hitting context window limits or long context performance degradation, and agents can perform very long trajectories that exceed a typical context window for long-running, complex tasks. A weaker version of this was previously possible with ad-hoc scaffolding and conversation summarization, but our first-class implementation, available via the Responses API, is integrated with the model and is highly performant.

225 

226How it works:

227 

2281. You use the Responses API as today, sending input items that include tool calls, user inputs, and assistant messages.

2292. When your context window grows large, you can invoke /compact to generate a new, compacted context window. Two things to note:

230 1. The context window that you send to /compact should fit within your model’s context window.

231 2. The endpoint is ZDR compatible and will return an “encrypted_content” item that you can pass into future requests.

2323. For subsequent calls to the /responses endpoint, you can pass your updated, compacted list of conversation items (including the added compaction item). The model retains key prior state with fewer conversation tokens.

233 

234For endpoint details see our `/responses/compact` [docs](https://developers.openai.com/api/reference/resources/responses/methods/compact).

235 

236### Tools

237 

2381. We strongly recommend using our exact `apply_patch` implementation as the model has been trained to excel at this diff format. For terminal commands we recommend our `shell` tool, and for plan/TODO items our `update_plan` tool should be most performant.

2392. If you prefer your agent to use more “terminal-like tools” (like `file_read()` instead of calling \`sed\` in the terminal), this model can reliably call them instead of terminal (following the instructions below)

2403. For other tools, including semantic search, MCPs, or other custom tools, they can work but it requires more tuning and experimentation.

241 

242#### Apply_patch

243 

244The easiest way to implement apply_patch is with our first-class implementation in the Responses API, but you can also use our freeform tool implementation with [context-free grammar](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_new_params_and_tools?utm_source=chatgpt.com#3-contextfree-grammar-cfg). Both are demonstrated below.

245 

246```python

247# Sample script to demonstrate the server-defined apply_patch tool

248 

249import json

250from pprint import pprint

251from typing import cast

252 

253from openai import OpenAI

254from openai.types.responses import ResponseInputParam, ToolParam

255 

256client = OpenAI()

257 

258## Shared tools and prompt

259user_request = """Add a cancel button that logs when clicked"""

260file_excerpt = """\

261export default function Page() {

262return (

263<div>

264 <p>Page component not implemented</p>

265 <button onClick={() => console.log("clicked")}>Click me</button>

266</div>

267);

268}

269"""

270 

271input_items: ResponseInputParam = [

272 {"role": "user", "content": user_request},

273 {

274 "type": "function_call",

275 "call_id": "call_read_file_1",

276 "name": "read_file",

277 "arguments": json.dumps({"path": ("/app/page.tsx")}),

278 },

279 {

280 "type": "function_call_output",

281 "call_id": "call_read_file_1",

282 "output": file_excerpt,

283 },

284]

285 

286read_file_tool: ToolParam = cast(

287 ToolParam,

288 {

289 "type": "function",

290 "name": "read_file",

291 "description": "Reads a file from disk",

292 "parameters": {

293 "type": "object",

294 "properties": {"path": {"type": "string"}},

295 "required": ["path"],

296 },

297 },

298)

299 

300### Get patch with built-in responses tool

301tools: list[ToolParam] = [

302 read_file_tool,

303 cast(ToolParam, {"type": "apply_patch"}),

304]

305 

306response = client.responses.create(

307 model="gpt-5.3-codex",

308 input=input_items,

309 tools=tools,

310 parallel_tool_calls=False,

311)

312 

313for item in response.output:

314 if item.type == "apply_patch_call":

315 print("Responses API apply_patch patch:")

316 pprint(item.operation)

317 # output:

318 # {'diff': '@@\n'

319 # ' return (\n'

320 # ' <div>\n'

321 # ' <p>Page component not implemented</p>\n'

322 # ' <button onClick={() => console.log("clicked")}>Click me</button>\n'

323 # '+ <button onClick={() => console.log("cancel clicked")}>Cancel</button>\n'

324 # ' </div>\n'

325 # ' );\n'

326 # ' }\n',

327 # 'path': '/app/page.tsx',

328 # 'type': 'update_file'}

329 

330### Get patch with custom tool implementation, including freeform tool definition and context-free grammar

331apply_patch_grammar = """

332start: begin_patch hunk+ end_patch

333begin_patch: "*** Begin Patch" LF

334end_patch: "*** End Patch" LF?

335 

336hunk: add_hunk | delete_hunk | update_hunk

337add_hunk: "*** Add File: " filename LF add_line+

338delete_hunk: "*** Delete File: " filename LF

339update_hunk: "*** Update File: " filename LF change_move? change?

340 

341filename: /(.+)/

342add_line: "+" /(.*)/ LF -> line

343 

344change_move: "*** Move to: " filename LF

345change: (change_context | change_line)+ eof_line?

346change_context: ("@@" | "@@ " /(.+)/) LF

347change_line: ("+" | "-" | " ") /(.*)/ LF

348eof_line: "*** End of File" LF

349 

350%import common.LF

351"""

352 

353tools_with_cfg: list[ToolParam] = [

354 read_file_tool,

355 cast(

356 ToolParam,

357 {

358 "type": "custom",

359 "name": "apply_patch_grammar",

360 "description": "Use the `apply_patch` tool to edit files. This is a FREEFORM tool, so do not wrap the patch in JSON.",

361 "format": {

362 "type": "grammar",

363 "syntax": "lark",

364 "definition": apply_patch_grammar,

365 },

366 },

367 ),

368]

369 

370response_cfg = client.responses.create(

371 model="gpt-5.3-codex",

372 input=input_items,

373 tools=tools_with_cfg,

374 parallel_tool_calls=False,

375)

376 

377for item in response_cfg.output:

378 if item.type == "custom_tool_call":

379 print("\n\nContext-free grammar apply_patch patch:")

380 print(item.input)

381 # Output

382 # *** Begin Patch

383 # *** Update File: /app/page.tsx

384 # @@

385 # <div>

386 # <p>Page component not implemented</p>

387 # <button onClick={() => console.log("clicked")}>Click me</button>

388 # + <button onClick={() => console.log("cancel clicked")}>Cancel</button>

389 # </div>

390 # );

391 # }

392 # *** End Patch

393```

394 

395```ruby

396require "openai"

397 

398client = OpenAI::Client.new

399 

400input = <<~PROMPT

401 Add a cancel button next to the save button in app/page.tsx.

402 Current file contents:

403 export default function Page() { return <button>Save</button>; }

404PROMPT

405response = client.responses.create(

406 model: "gpt-5.3-codex", input: input,

407 tools: [{ type: :apply_patch }], parallel_tool_calls: false

408)

409response.output.each do |item|

410 pp(item.operation) if item.is_a?(OpenAI::Responses::ResponseApplyPatchToolCall)

411end

412 

413APPLY_PATCH_GRAMMAR = <<~GRAMMAR

414 

415 start: begin_patch hunk+ end_patch

416 begin_patch: "*** Begin Patch" LF

417 end_patch: "*** End Patch" LF?

418 

419 hunk: add_hunk | delete_hunk | update_hunk

420 add_hunk: "*** Add File: " filename LF add_line+

421 delete_hunk: "*** Delete File: " filename LF

422 update_hunk: "*** Update File: " filename LF change_move? change?

423 

424 filename: /(.+)/

425 add_line: "+" /(.*)/ LF -> line

426 

427 change_move: "*** Move to: " filename LF

428 change: (change_context | change_line)+ eof_line?

429 change_context: ("@@" | "@@ " /(.+)/) LF

430 change_line: ("+" | "-" | " ") /(.*)/ LF

431 eof_line: "*** End of File" LF

432 

433 %import common.LF

434 

435GRAMMAR

436 

437response = client.responses.create(

438 model: "gpt-5.3-codex", input: input,

439 tools: [

440 {

441 type: :custom,

442 name: "apply_patch",

443 description: "Apply a patch to update files.",

444 format: {

445 type: :grammar,

446 syntax: :lark,

447 definition: APPLY_PATCH_GRAMMAR

448 }

449 }

450 ],

451 parallel_tool_calls: false

452)

453response.output.each do |item|

454 puts(item.input) if item.is_a?(OpenAI::Responses::ResponseCustomToolCall)

455end

456```

457 

458 

459Patches objects the Responses API tool can be implemented by following this [example](https://github.com/openai/openai-agents-python/blob/main/examples/tools/apply_patch.py) and patches from the freeform tool can be applied with the logic in our canonical GPT-5 [apply_patch.py](https://github.com/openai/openai-cookbook/blob/main/examples/gpt-5/apply_patch.py%20) implementation.

460 

461#### Shell_command

462 

463This is our default shell tool. Note that we have seen better performance with a command type “string” rather than a list of commands.

464 

465```json

466{

467 "type": "function",

468 "function": {

469 "name": "shell_command",

470 "description": "Runs a shell command and returns its output.\n- Always set the `workdir` param when using the shell_command function. Do not use `cd` unless absolutely necessary.",

471 "strict": false,

472 "parameters": {

473 "type": "object",

474 "properties": {

475 "command": {

476 "type": "string",

477 "description": "The shell script to execute in the user's default shell"

478 },

479 "workdir": {

480 "type": "string",

481 "description": "The working directory to execute the command in"

482 },

483 "timeout_ms": {

484 "type": "number",

485 "description": "The timeout for the command in milliseconds"

486 },

487 "with_escalated_permissions": {

488 "type": "boolean",

489 "description": "Whether to request escalated permissions. Set to true if command needs to be run without sandbox restrictions"

490 },

491 "justification": {

492 "type": "string",

493 "description": "Only set if with_escalated_permissions is true. 1-sentence explanation of why we want to run this command."

494 }

495 },

496 "required": ["command"],

497 "additionalProperties": false

498 }

499 }

500}

501```

502 

503If you’re using Windows PowerShell, update to this tool description.

504 

505```text

506Runs a shell command and returns its output. The arguments you pass will be invoked via PowerShell (e.g., ["pwsh", "-NoLogo", "-NoProfile", "-Command", "<cmd>"]). Always fill in workdir; avoid using cd in the command string.

507```

508 

509You can check out codex-cli for the implementation for `exec_command`, which launches a long-lived PTY when you need streaming output, REPLs, or interactive sessions; and `write_stdin`, to feed extra keystrokes (or just poll output) for an existing exec_command session.

510 

511#### Update Plan

512 

513This is our default TODO tool; feel free to customize as you’d prefer. See the `## Plan tool` section of our starter prompt for additional instructions to maintain hygiene and tweak behavior.

514 

515```json

516{

517 "type": "function",

518 "function": {

519 "name": "update_plan",

520 "description": "Updates the task plan.\nProvide an optional explanation and a list of plan items, each with a step and status.\nAt most one step can be in_progress at a time.",

521 "strict": false,

522 "parameters": {

523 "type": "object",

524 "properties": {

525 "explanation": {

526 "type": "string"

527 },

528 "plan": {

529 "type": "array",

530 "items": {

531 "type": "object",

532 "properties": {

533 "step": {

534 "type": "string"

535 },

536 "status": {

537 "type": "string",

538 "description": "One of: pending, in_progress, completed"

539 }

540 },

541 "additionalProperties": false,

542 "required": ["step", "status"]

543 },

544 "description": "The list of steps"

545 }

546 },

547 "additionalProperties": false,

548 "required": ["plan"]

549 }

550 }

551}

552```

553 

554#### View_image

555 

556This is a basic function used in codex-cli for the model to view images.

557 

558```json

559{

560 "type": "function",

561 "function": {

562 "name": "view_image",

563 "description": "Attach a local image (by filesystem path) to the conversation context for this turn.",

564 "strict": false,

565 "parameters": {

566 "type": "object",

567 "properties": {

568 "path": {

569 "type": "string",

570 "description": "Local filesystem path to an image file"

571 }

572 },

573 "additionalProperties": false,

574 "required": ["path"]

575 }

576 }

577}

578```

579 

580### Dedicated terminal-wrapping tools

581 

582If you would prefer your codex agent to use terminal-wrapping tools (like a dedicated `list_dir(‘.’)` tool instead of `terminal(‘ls .’)`, this generally works well. We see the best results when the name of the tool, the arguments, and the output are as close as possible to those from the underlying command, so it’s as in-distribution as possible for the model (which was primarily trained using a dedicated terminal tool). For example, if you notice the model using git via the terminal and would prefer it to use a dedicated tool, we found that creating a related tool, and adding a directive in the prompt to only use that tool for git commands, fully mitigated the model’s terminal usage for git commands.

583 

584```python

585GIT_TOOL = {

586 "type": "function",

587 "name": "git",

588 "description": (

589 "Execute a git command in the repository root. Behaves like running git in the"

590 " terminal; supports any subcommand and flags. The command can be provided as a"

591 " full git invocation (e.g., `git status -sb`) or just the arguments after git"

592 " (e.g., `status -sb`)."

593 ),

594 "parameters": {

595 "type": "object",

596 "properties": {

597 "command": {

598 "type": "string",

599 "description": (

600 "The git command to execute. Accepts either a full git invocation or"

601 " only the subcommand/args."

602 ),

603 },

604 "timeout_sec": {

605 "type": "integer",

606 "minimum": 1,

607 "maximum": 1800,

608 "description": "Optional timeout in seconds for the git command.",

609 },

610 },

611 "required": ["command"],

612 },

613}

614 

615TOOLS = [GIT_TOOL]

616 

617PROMPT_TOOL_USE_DIRECTIVE = (

618 "- Strictly avoid raw `cmd`/terminal for Git operations. Use the dedicated "

619 "`git` tool instead."

620)

621```

622 

623```ruby

624require "json"

625 

626GIT_TOOL = {

627 "type" => "function",

628 "name" => "git",

629 "description" => "Execute a git command in the repository root. Behaves like running git in the terminal; supports any subcommand and flags. The command can be provided as a full git invocation (e.g., `git status -sb`) or just the arguments after git (e.g., `status -sb`).",

630 "parameters" => {

631 "type" => "object",

632 "properties" => {

633 "command" => {

634 "type" => "string",

635 "description" => "The git command to execute. Accepts either a full git invocation or only the subcommand/args."

636 },

637 "timeout_sec" => {

638 "type" => "integer",

639 "minimum" => 1,

640 "maximum" => 1800,

641 "description" => "Optional timeout in seconds for the git command."

642 }

643 },

644 "required" => ["command"]

645 }

646}

647 

648TOOLS = [GIT_TOOL]

649PROMPT_TOOL_USE_DIRECTIVE = "- Strictly avoid raw `cmd`/terminal for Git operations. Use the dedicated `git` tool instead."

650puts(JSON.generate(TOOLS))

651puts(PROMPT_TOOL_USE_DIRECTIVE)

652```

653 

654 

655### Other Custom Tools (web search, semantic search, memory, etc.)

656 

657The model hasn’t necessarily been post-trained to excel at these tools, but we have seen success here as well. To get the most out of these tools, we recommend:

658 

6591. Making the tool names and arguments as semantically “correct” as possible, for example “search” is ambiguous but “semantic_search” clearly indicates what the tool does, relative to other potential search-related tools you might have. “Query” would be a good param name for this tool.

6602. Be explicit in your prompt about when, why, and how to use these tools, including good and bad examples.

6613. It could also be helpful to make the results look different from outputs the model is accustomed to seeing from other tools, for example ripgrep results should look different from semantic search results to avoid the model collapsing into old habits.

662 

663### Parallel Tool Calling

664 

665In codex-cli, when parallel tool calling is enabled, the responses API request sets `parallel_tool_calls: true` and the following snippet is added to the system instructions:

666 

667```text

668## Exploration and reading files

669 

670- **Think first.** Before any tool call, decide ALL files/resources you will need.

671- **Batch everything.** If you need multiple files (even from different places), read them together.

672- **multi_tool_use.parallel** Use `multi_tool_use.parallel` to parallelize tool calls and only this.

673- **Only make sequential calls if you truly cannot know the next file without seeing a result first.**

674- **Workflow:** (a) plan all needed reads → (b) issue one parallel batch → (c) analyze results → (d) repeat if new, unpredictable reads arise.

675 

676**Additional notes**:

677- Always maximize parallelism. Never read files one-by-one unless logically unavoidable.

678- This concerns every read/list/search operations including, but not only, `cat`, `rg`, `sed`, `ls`, `git show`, `nl`, `wc`, ...

679- Do not try to parallelize using scripting or anything else than `multi_tool_use.parallel`.

680```

681 

682We've found it to be helpful and more in-distribution if parallel tool call items and responses are ordered in the following way:

683 

684```text

685function_call

686function_call

687function_call_output

688function_call_output

689```

690 

691### Tool Response Truncation

692 

693We recommend doing tool call response truncation as follows to be as in-distribution for the model as possible:

694 

695- Limit to 10k tokens. You can cheaply approximate this by computing `num_bytes/4`.

696- If you hit the truncation limit, you should use half of the budget for the beginning, half for the end, and truncate in the middle with `…3 tokens truncated…`

697 

698### New features in GPT-5.3 Codex

699 

700#### Preamble messages

701 

702The Responses API includes a `phase` parameter intended to prevent early stopping and other misbehavior when preamble messages are requested by the prompt. Correctly implementing this parameter is required for `gpt-5.3-codex`; otherwise, significant performance degradation can occur.

703 

704#### Phase

705 

706To better support preamble messages with `gpt-5.3-codex`, the Responses API includes a `phase` field designed to prevent early stopping on longer-running tasks and other misbehaviors.

707 

708##### Values

709 

710`phase` is one of:

711 

712- `null`

713- `"commentary"`

714- `"final_answer"`

715 

716##### Where it appears

717 

718You’ll receive `phase` on assistant output items (for example, `output_item.done`). Your integration must persist assistant output items, including their `phase`, and pass those assistant items back in subsequent requests.

719 

720**Important:** `phase` is only supported on assistant items. Do not add `phase` to user messages.

721 

722##### How it’s used downstream

723 

724When the model marks an output item with:

725 

726- `phase: "commentary"`: the corresponding assistant message should be treated as commentary/preamble-style content.

727- `phase: "final_answer"`: the corresponding assistant message should be treated as the final closeout.

728 

729Correctly preserving `phase` on assistant items is required for `gpt-5.3-codex`. If assistant `phase` metadata is dropped during history reconstruction, significant performance degradation can occur.

730 

731#### Preambles & Personality

732 

733Preambles are messages sent along with tool calls that provide user updates while working: short, human-readable progress and intent snapshots that keep the user oriented without turning the transcript into a tool-call log. GPT-5.3-Codex preambles have been tuned toward the following characteristics:

734 

735- Acknowledge then plan before any tool calls (1 sentence acknowledgement, 1–2 sentence plan).

736- Keep most updates to 1–2 sentences, and use longer updates only at real milestones.

737- Cadence: aim every 1–3 execution steps; hard floor: at least within every 6 steps or 10 tool calls.

738- Content per update: outcome/impact so far, next 1–3 steps, and open questions/learnings when present.

739- Tone: real person pairing, low-ceremony; avoid headings/status labels and log voice.

740 

741##### Personality (Friendly vs Pragmatic)

742 

743Personality is the higher-level vibe and collaboration posture that sits above preamble mechanics (cadence, length, and grounding). It affects word choice, how eagerly the model explains tradeoffs, and how much warmth it brings to the interaction.

744 

745The Codex app and CLI ship with support for two personalities provided here as example implementations for your harness.

746 

747###### Friendly

748 

749- More human, partner-y pairing energy.

750- Slightly more acknowledgement, reassurance, and context-setting.

751- Better when the user benefits from narrative orientation (onboarding, ambiguous tasks, higher-stakes changes).

752 

753###### Example Friendly personality prompt snippet from codex-cli

754 

755This snippet can be used in your system prompt to steer the pair programming personality of the model.

756 

757```text

758# Personality

759 

760You optimize for team morale and being a supportive teammate as much as code quality. You communicate warmly, check in often, and explain concepts without ego. You excel at pairing, onboarding, and unblocking others. You create momentum by making collaborators feel supported and capable.

761 

762## Values

763You are guided by these core values:

764* Empathy: Interprets empathy as meeting people where they are - adjusting explanations, pacing, and tone to maximize understanding and confidence.

765* Collaboration: Sees collaboration as an active skill: inviting input, synthesizing perspectives, and making others successful.

766* Ownership: Takes responsibility not just for code, but for whether teammates are unblocked and progress continues.

767 

768## Tone & User Experience

769Your voice is warm, encouraging, and conversational. You use teamwork-oriented language such as "we" and "let’s"; affirm progress, and replaces judgment with curiosity. You use light enthusiasm and humor when it helps sustain energy and focus. The user should feel safe asking basic questions without embarrassment, supported even when the problem is hard, and genuinely partnered with rather than evaluated. Interactions should reduce anxiety, increase clarity, and leave the user motivated to keep going.

770 

771You are NEVER curt or dismissive.

772 

773You are a patient and enjoyable collaborator: unflappable when others might get frustrated, while being an enjoyable, easy-going personality to work with. Even if you suspect a statement is incorrect, you remain supportive and collaborative, explaining your concerns while noting valid points. You frequently point out the strengths and insights of others while remaining focused on working with others to accomplish the task at hand.

774 

775## Escalation

776You escalate gently and deliberately when decisions have non-obvious consequences or hidden risk. Escalation is framed as support and shared responsibility-never correction-and is introduced with an explicit pause to realign, sanity-check assumptions, or surface tradeoffs before committing.

777```

778 

779###### Pragmatic

780 

781- More terse, direct, let’s ship delivery.

782- Fewer social flourishes; higher ratio of actionable information per token.

783- Better when latency/throughput matters, or your users already know the workflow and just want progress and results.

784 

785#### Troubleshooting & Metaprompting

786 

787Common failure modes we’ve been explicitly tracking:

788 

789- Overthinking / long time before first useful action (tool call or concrete plan).

790- Loggy / unnatural status updates instead of pair programmer collaboration.

791- Awkward preamble phrasing and repetitive tics ("Good catch", "Aha", "Got it–", etc.).

792 

793##### Metaprompting for targeted fixes

794 

795Failure modes like the ones above can typically be addressed through metaprompting. It’s possible to ask the model at the end of a turn that didn’t perform up to expectations how to improve its own instructions. The following prompt was used to produce some of the solutions to overthinking problems above and can be modified to meet your particular needs.

796 

797```text

798That was a high quality response, thanks! It seemed like it took you a while to finish responding though. Is there a way to clarify your instructions so you can get to a response as good as this faster next time? It’s extremely important to be efficient when providing these responses or users won’t get the most out of them in time. Let’s see if we can improve!

799think through the response you gave above

800read through your instructions starting from "" and look for anything that might have made you take longer to formulate a high quality response than you needed

801write out targeted (but generalized) additions/changes/deletions to your instructions to make a request like this one faster next time with the same level of quality

802```

803 

804When metaprompting inside a specific context, it is important to generate responses a few times if possible and pay attention to elements of the responses that are common between them. Some improvements or changes the model proposes might be overly specific to that particular situation, but you can often simplify them to arrive at a general improvement. We recommend creating an eval to measure whether a particular prompt change is better or worse for your particular use case.

805 

806##### Some examples

807 

808- For overthinking / slow starts: ask it to propose instruction changes that reduce time-to-first-tool-call or first concrete plan.

809- For overly loggy preambles: ask it to rewrite your user updates instructions to satisfy your particular preference constraints.

810 

guides/latest-model/gpt-5.4.md +0 −1383 deleted

File Deleted View Diff

1# Using GPT-5.4

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7[GPT-5.4](https://developers.openai.com/api/docs/models/gpt-5.4) was released as a frontier model for professional work across the API and Codex. It helps developers analyze complex information, build production software, and automate multi-step workflows.

8 

9Within the GPT-5.4 generation, `gpt-5.4` is the general-purpose model for workflows that move between software engineering, reasoning, writing, and tool use.

10 

11This guide covers key features of the GPT-5 model family and how to get the most out of GPT-5.4.

12 

13## What's new

14 

15Compared with the previous GPT-5.2 model, GPT-5.4 shows improvements in:

16 

17- Coding, document understanding, tool use, and instruction following

18- Image perception and multimodal tasks

19- Long-running task execution and multi-step agent workflows

20- Token efficiency and end-to-end performance on tool-heavy workloads

21- Web search and multi-source synthesis for hard-to-locate information

22- Document-heavy and spreadsheet-heavy business workflows in customer service, analytics, and finance

23 

24GPT-5.4 brings the coding capabilities of GPT-5.3-Codex to our flagship frontier model. Developers can generate production-quality code, build polished front-end UI, follow repo-specific patterns, and handle multi-file changes with fewer retries. It also has a strong out-of-the-box coding personality, so teams spend less time on prompt tuning.

25 

26For agentic workloads, GPT-5.4 reduces end-to-end time across multi-step trajectories and often completes tasks with fewer tokens and tool calls. This makes agents more responsive and lowers the cost of operating complex workflows at scale in the API and Codex.

27 

28### New features in GPT-5.4

29 

30Like earlier GPT-5 models, GPT-5.4 supports custom tools, parameters to control verbosity and reasoning, and an allowed tools list. GPT-5.4 also introduces several capabilities that make it easier to build powerful agent systems, operate over larger bodies of information, and run more reliable automated workflows:

31 

32- **`tool_search` in the API:** GPT-5.4 improves tool search for larger tool ecosystems by using deferred tool loading. This makes tools searchable, loads only the relevant definitions, reduces token usage, and improves tool selection accuracy in real deployments. Learn more in the [tool search guide](https://developers.openai.com/api/docs/guides/tools-tool-search).

33- **1M token context window:** GPT-5.4 supports up to a 1M token context window, making it easier to analyze entire codebases, long document collections, or extended agent trajectories in a single request. Read more in the [1M context window](#1m-context-window) section.

34- **Built-in computer use:** GPT-5.4 is the first mainline model with built-in computer-use capabilities, enabling agents to interact directly with software to complete, verify, and fix tasks in a build-run-verify-fix loop. Learn more in the [computer use guide](https://developers.openai.com/api/docs/guides/tools-computer-use).

35- **Native compaction support:** GPT-5.4 is the first mainline model trained to support compaction, enabling longer agent trajectories while preserving key context.

36 

37## Model, API, and feature updates

38 

39Within this model generation, `gpt-5.4` is the general-purpose model for both broad tasks and coding. For more difficult problems, `gpt-5.4-pro` uses more compute to think longer and provide more consistent answers.

40 

41For smaller, faster variants, start with `gpt-5.4-mini` or `gpt-5.4-nano`.

42 

43To help you pick the model that best fits your use case, consider these tradeoffs:

44 

45| Variant | Best for |

46| ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |

47| [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) | General-purpose work, including complex reasoning, broad world knowledge, and code-heavy or multi-step agentic tasks |

48| [`gpt-5.4-pro`](https://developers.openai.com/api/docs/models/gpt-5.4-pro) | Tough problems that may take longer to solve and need deeper reasoning |

49| [`gpt-5.4-mini`](https://developers.openai.com/api/docs/models/gpt-5.4-mini) | High-volume coding, computer use, and agent workflows that still need strong reasoning |

50| [`gpt-5.4-nano`](https://developers.openai.com/api/docs/models/gpt-5.4-nano) | High-throughput tasks where speed and cost matter most |

51 

52### Lower reasoning effort

53 

54The `reasoning.effort` parameter controls how many reasoning tokens the model generates before producing a response. Earlier reasoning models like o3 supported only `low`, `medium`, and `high`: `low` favored speed and fewer tokens, while `high` favored more thorough reasoning.

55 

56GPT-5.2 and GPT-5.4 support `none` as their lowest reasoning effort for lower-latency interactions. It is the default setting for both models. If you need more thinking, slowly increase to `medium` and experiment with results.

57 

58With reasoning effort set to `none`, prompting is important. To improve the model's reasoning quality, even with the default settings, encourage it to “think” or outline its steps before answering.

59 

60Reasoning effort set to none

61 

62```javascript

63import OpenAI from "openai";

64const openai = new OpenAI();

65 

66const response = await openai.responses.create({

67 model: "gpt-5.4",

68 input:

69 "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

70 reasoning: {

71 effort: "none",

72 },

73});

74 

75console.log(response);

76```

77 

78```python

79from openai import OpenAI

80 

81client = OpenAI()

82 

83response = client.responses.create(

84 model="gpt-5.4",

85 input="Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

86 reasoning={"effort": "none"},

87)

88 

89print(response)

90```

91 

92```go

93package main

94 

95import (

96 "context"

97 "fmt"

98 

99 "github.com/openai/openai-go/v3"

100 "github.com/openai/openai-go/v3/responses"

101 "github.com/openai/openai-go/v3/shared"

102)

103 

104func main() {

105 client := openai.NewClient()

106 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{

107 Model: "gpt-5.4",

108 Input: responses.ResponseNewParamsInputUnion{OfString: openai.String("Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?")},

109 Reasoning: shared.ReasoningParam{Effort: shared.ReasoningEffortNone},

110 })

111 if err != nil {

112 panic(err)

113 }

114 fmt.Println(response)

115}

116```

117 

118```java

119import com.openai.client.OpenAIClient;

120import com.openai.client.okhttp.OpenAIOkHttpClient;

121import com.openai.models.Reasoning;

122import com.openai.models.ReasoningEffort;

123import com.openai.models.responses.ResponseCreateParams;

124 

125ResponseCreateParams params =

126 ResponseCreateParams.builder()

127 .model("gpt-5.4")

128 .input("Explain the bug and propose a fix.")

129 .reasoning(Reasoning.builder().effort(ReasoningEffort.NONE).build())

130 .build();

131 

132client.responses().create(params).output().stream()

133 .flatMap(item -> item.message().stream())

134 .flatMap(message -> message.content().stream())

135 .flatMap(content -> content.outputText().stream())

136 .forEach(text -> System.out.println(text.text()));

137```

138 

139```csharp

140using OpenAI.Responses;

141#pragma warning disable OPENAI001

142 

143string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;

144ResponsesClient client = new(key);

145 

146CreateResponseOptions options = new()

147{

148 Model = "gpt-5.4",

149 ReasoningOptions = new ResponseReasoningOptions

150 {

151 ReasoningEffortLevel = ResponseReasoningEffortLevel.None,

152 },

153};

154options.InputItems.Add(

155 ResponseItem.CreateUserMessageItem(

156 "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?"

157 )

158);

159 

160ResponseResult response = await client.CreateResponseAsync(options);

161Console.WriteLine(response.GetOutputText());

162```

163 

164```ruby

165require "openai"

166 

167client = OpenAI::Client.new

168response = client.responses.create(

169 model: "gpt-5.4",

170 reasoning: { effort: :minimal },

171 input: "Explain the bug and propose a fix."

172)

173puts(response.output_text)

174```

175 

176```bash

177curl --request POST \

178 --url https://api.openai.com/v1/responses \

179 --header "Authorization: Bearer $OPENAI_API_KEY" \

180 --header 'Content-type: application/json' \

181 --data '{

182 "model": "gpt-5.4",

183 "input": "Think carefully and outline your steps before answering. How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

184 "reasoning": {

185 "effort": "none"

186 }

187}'

188```

189 

190 

191### Verbosity

192 

193Verbosity determines how many output tokens are generated. Lowering the number of tokens reduces overall latency. While the model's reasoning approach stays mostly the same, the model finds ways to answer more concisely—which can either improve or diminish answer quality, depending on your use case. Here are some scenarios for both ends of the verbosity spectrum:

194 

195- **High verbosity:** Use when you need the model to provide thorough explanations of documents or perform extensive code refactoring.

196- **Low verbosity:** Best for situations where you want concise answers or focused code generation, such as SQL queries.

197 

198GPT-5 made this option configurable as one of `high`, `medium`, or `low`. With GPT-5.4, verbosity remains configurable and defaults to `medium`.

199 

200When generating code with GPT-5.4, `medium` and `high` verbosity levels yield longer, more structured code with inline explanations, while `low` verbosity produces shorter, more concise code with minimal commentary.

201 

202Control verbosity

203 

204```javascript

205import OpenAI from "openai";

206const openai = new OpenAI();

207 

208const response = await openai.responses.create({

209 model: "gpt-5.4",

210 input:

211 "What is the answer to the ultimate question of life, the universe, and everything?",

212 text: {

213 verbosity: "low",

214 },

215});

216 

217console.log(response);

218```

219 

220```python

221from openai import OpenAI

222 

223client = OpenAI()

224 

225response = client.responses.create(

226 model="gpt-5.4",

227 input="What is the answer to the ultimate question of life, the universe, and everything?",

228 text={"verbosity": "low"},

229)

230 

231print(response)

232```

233 

234```go

235package main

236 

237import (

238 "context"

239 "fmt"

240 

241 "github.com/openai/openai-go/v3"

242 "github.com/openai/openai-go/v3/responses"

243)

244 

245func main() {

246 client := openai.NewClient()

247 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{

248 Model: "gpt-5.4",

249 Input: responses.ResponseNewParamsInputUnion{OfString: openai.String("What is the answer to the ultimate question of life, the universe, and everything?")},

250 Text: responses.ResponseTextConfigParam{Verbosity: responses.ResponseTextConfigVerbosityLow},

251 })

252 if err != nil {

253 panic(err)

254 }

255 fmt.Println(response)

256}

257```

258 

259```java

260import com.openai.client.OpenAIClient;

261import com.openai.client.okhttp.OpenAIOkHttpClient;

262import com.openai.models.responses.ResponseCreateParams;

263import com.openai.models.responses.ResponseTextConfig;

264 

265ResponseCreateParams params =

266 ResponseCreateParams.builder()

267 .model("gpt-5.4")

268 .input("Explain the bug and propose a fix.")

269 .text(ResponseTextConfig.builder().verbosity(ResponseTextConfig.Verbosity.LOW).build())

270 .build();

271 

272client.responses().create(params).output().stream()

273 .flatMap(item -> item.message().stream())

274 .flatMap(message -> message.content().stream())

275 .flatMap(content -> content.outputText().stream())

276 .forEach(text -> System.out.println(text.text()));

277```

278 

279```ruby

280require "openai"

281 

282client = OpenAI::Client.new

283response = client.responses.create(

284 model: "gpt-5.4",

285 text: { verbosity: :low },

286 input: "Explain the bug and propose a fix."

287)

288puts(response.output_text)

289```

290 

291```bash

292curl --request POST \

293 --url https://api.openai.com/v1/responses \

294 --header "Authorization: Bearer $OPENAI_API_KEY" \

295 --header 'Content-type: application/json' \

296 --data '{

297 "model": "gpt-5.4",

298 "input": "What is the answer to the ultimate question of life, the universe, and everything?",

299 "text": {

300 "verbosity": "low"

301 }

302}'

303```

304 

305 

306You can still steer verbosity through prompting after setting it to `low` in the API. The verbosity parameter defines a general token range at the system prompt level, but the actual output is flexible to both developer and user prompts within that range.

307 

308#### 1M context window

309 

3101M token context window was introduced with GPT-5.4, making it easier to analyze entire codebases, long document collections, or extended agent trajectories in a single request.

311 

312We have separate standard pricing for requests under 272K and over 272K tokens, available in the [pricing docs](https://developers.openai.com/api/docs/pricing). If you use [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode), any prompt above 272K tokens is automatically processed at standard rates.

313 

314Long context pricing stacks with other pricing modifiers such as data residency and batch.

315 

316We have different rate limits for requests under 272K tokens and over 272K tokens; this is available on the [GPT-5.4 model page](https://developers.openai.com/api/docs/models/gpt-5.4).

317 

318## Using tools with GPT-5.4

319 

320GPT-5.4 has been post-trained on specific tools. See the [tools docs](https://developers.openai.com/api/docs/guides/tools) for more specific guidance.

321 

322### Computer use tool

323 

324Computer use lets GPT-5.4 operate software through the user interface by inspecting screenshots and returning structured actions for your harness to execute. It is a good fit for browser or desktop workflows where a person could complete the task through the UI, such as navigating a site, filling out forms, or validating that a change actually worked.

325 

326Use it in an isolated browser or VM, and keep a human in the loop for high-impact actions. The full guide covers the built-in Responses API loop, custom harness patterns, and code-execution-based setups.

327 

328[Computer use guide

329 

330 

331 

332 Learn how to run the built-in computer tool safely and integrate it with

333 your own harness.](https://developers.openai.com/api/docs/guides/tools-computer-use)

334 

335### Tool search tool

336 

337Tool search lets GPT-5.4 defer large tool surfaces until runtime so the model loads only the definitions it needs. This is most useful when you have many functions, `namespaces`, or MCP tools and want to reduce token usage, preserve cache performance, and improve latency without exposing every schema up front.

338 

339Use hosted tool search when the candidate tools are already known at request time, or client-executed tool search when your application needs to decide what to load dynamically. The full guide also covers best practices for `namespaces`, MCP servers, and deferred loading.

340 

341[Tool search guide

342 

343 

344 

345 Learn how to defer tool definitions and load the right subset at runtime.](https://developers.openai.com/api/docs/guides/tools-tool-search)

346 

347### Custom tools

348 

349When the GPT-5 model family launched, we introduced a new capability called custom tools, which lets models send any raw text as tool call input but still constrain outputs if desired. This tool behavior remains true in GPT-5.4.

350 

351[Function calling guide

352 

353 

354 

355 Learn about custom tools in the function calling guide.](https://developers.openai.com/api/docs/guides/function-calling)

356 

357#### Freeform inputs

358 

359Define your tool with `type: custom` to enable models to send plaintext inputs directly to your tools, rather than being limited to structured JSON. The model can send any raw text—code, SQL queries, shell commands, configuration files, or long-form prose—directly to your tool.

360 

361```json

362{

363 "type": "custom",

364 "name": "code_exec",

365 "description": "Executes arbitrary python code"

366}

367```

368 

369#### Constraining outputs

370 

371GPT-5.4 supports context-free grammars (`CFGs`) for custom tools, letting you provide a Lark grammar to constrain outputs to a specific syntax or DSL. Attaching a CFG, for example a SQL or DSL grammar, ensures the assistant's text matches your grammar.

372 

373This enables precise, constrained tool calls or structured responses and lets you enforce strict syntactic or domain-specific formats directly in GPT-5.4's function calling, improving control and reliability for complex or constrained domains.

374 

375#### Best practices for custom tools

376 

377- **Write concise, explicit tool descriptions.** The model chooses what to send based on your description; state explicitly if you want it to always call the tool.

378- **Validate outputs on the server side**. Freeform strings are powerful but require safeguards against injection or unsafe commands.

379 

380### Allowed tools

381 

382The `allowed_tools` parameter under `tool_choice` lets you pass N tool definitions but restrict the model to only M (&lt; N) of them. List your full toolkit in `tools`, and then use an `allowed_tools` block to name the subset and specify a mode—either `auto` (the model may pick any of those) or `required` (the model must invoke one).

383 

384[Function calling guide

385 

386 

387 

388 Learn about the allowed tools option in the function calling guide.](https://developers.openai.com/api/docs/guides/function-calling)

389 

390By separating all possible tools from the subset that can be used _now_, you gain greater safety, predictability, and improved prompt caching. You also avoid brittle prompt engineering, such as hard-coded call order. GPT-5.4 dynamically invokes or requires specific functions mid-conversation while reducing the risk of unintended tool usage over long contexts.

391 

392| | **Standard Tools** | **Allowed Tools** |

393| ---------------- | ----------------------------------------- | ------------------------------------------------------------- |

394| Model's universe | All tools listed under **`"tools": […]`** | Only the subset under **`"tools": […]`** in **`tool_choice`** |

395| Tool invocation | Model may or may not call any tool | Model restricted to (or required to call) chosen tools |

396| Purpose | Declare available capabilities | Constrain which capabilities are actually used |

397 

398```json

399{

400 "tool_choice": {

401 "type": "allowed_tools",

402 "mode": "auto",

403 "tools": [

404 { "type": "function", "name": "get_weather" },

405 { "type": "function", "name": "search_docs" }

406 ]

407 }

408}

409```

410 

411For a more detailed overview of all of these new features, see the [prompt guidance for GPT-5.4](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.4#prompting-best-practices).

412 

413### Preambles

414 

415Preambles are brief, user-visible explanations that GPT-5.4 generates before invoking any tool or function, outlining its intent or plan—for example, “why I'm calling this tool.” They appear after the chain of thought and before the actual tool call, making the model's reasoning easier to understand and debug while supporting precise steering.

416 

417By letting GPT-5.4 “think out loud” before each tool call, preambles boost tool-calling accuracy (and overall task success) without bloating reasoning overhead. To enable preambles, add a system or developer instruction—for example: “Before you call a tool, explain why you are calling it.” GPT-5.4 adds a concise rationale to each specified tool call. The model may also output multiple messages between tool calls, which can enhance the interaction experience—particularly for minimal reasoning or latency-sensitive use cases.

418 

419For more on using preambles, see the [GPT-5 prompting cookbook](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_prompting_guide#tool-preambles).

420 

421## Migration quickstart

422 

423GPT-5.4 works best with the Responses API, which supports preserving reasoning context between turns to improve performance. Read below to migrate from your current model or API.

424 

425### Migrating from other models to GPT-5.4

426 

427Use the [OpenAI Docs

428 skill](https://github.com/openai/skills/tree/main/skills/.system/openai-docs)

429 when migrating existing prompts or workflows to GPT-5.4. It's available in our

430 public skills repository and the Codex desktop app.

431 

432While the model should be close to a drop-in replacement for GPT-5.2, there are a few key changes to call out. See [Prompt guidance for GPT-5.4](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.4#prompting-best-practices) for specific updates to make in your prompts.

433 

434Using GPT-5 models with the Responses API provides improved intelligence because of the API design. The Responses API can pass the previous turn's CoT to the model. This leads to fewer generated reasoning tokens, higher cache hit rates, and less latency. To learn more, see an [in-depth guide](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items) on the benefits of the Responses API.

435 

436When migrating to GPT-5.4 from an older OpenAI model, start by experimenting with reasoning levels and prompting strategies. Use the [prompt optimizer](https://platform.openai.com/chat/edit?models=gpt-5.4&optimize=true) to update your prompts for GPT-5.4 based on current best practices, then follow this model-specific guidance:

437 

438- **`gpt-5.2`**: `gpt-5.4` with default settings is meant to be a drop-in replacement.

439- **o3**: `gpt-5.4` with `medium` or `high` reasoning. Start with `medium` reasoning with prompt tuning, then increase to `high` if you aren't getting the results you want.

440- **`gpt-4.1`**: `gpt-5.4` with `none` reasoning. Start with `none` and tune your prompts; increase if you need better performance.

441- **`o4-mini` or `gpt-4.1-mini`**: `gpt-5.4-mini` with prompt tuning is a great replacement.

442- **`gpt-4.1-nano`**: `gpt-5.4-nano` with prompt tuning is a great replacement.

443 

444### New `phase` parameter

445 

446For long-running or tool-heavy GPT-5.4 flows in the Responses API, use the assistant message `phase` field to avoid early stopping and other misbehavior.

447 

448`phase` is optional at the API level, but we highly recommend using it. Use `phase: "commentary"` for intermediate assistant updates (such as preambles before tool calls) and `phase: "final_answer"` for the completed answer. Do not add `phase` to user messages.

449 

450If you use `previous_response_id`, that is usually the simplest path because

451 prior assistant state is preserved. If you replay assistant history manually,

452 preserve each original `phase` value.

453 

454Missing or dropped `phase` can cause preambles to be treated as final answers

455in those workflows. For additional guidance and examples, see the [GPT-5.4

456prompting guide](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.4#phase-parameter).

457 

458Round-trip assistant phase values

459 

460```javascript

461import OpenAI from "openai";

462const client = new OpenAI();

463 

464const response = await client.responses.create({

465 model: "gpt-5.4",

466 input: [

467 {

468 role: "assistant",

469 phase: "commentary",

470 content:

471 "I’ll inspect the logs and then summarize root cause and remediation.",

472 },

473 {

474 role: "assistant",

475 phase: "final_answer",

476 content: "Root cause: cache invalidation race.",

477 },

478 {

479 role: "user",

480 content: "Great—now give me a rollout-safe fix plan.",

481 },

482 ],

483});

484 

485console.log(response.output_text);

486```

487 

488```python

489from openai import OpenAI

490 

491client = OpenAI()

492 

493response = client.responses.create(

494 model="gpt-5.4",

495 input=[

496 {

497 "role": "assistant",

498 "phase": "commentary",

499 "content": "I’ll inspect the logs and then summarize root cause and remediation.",

500 },

501 {

502 "role": "assistant",

503 "phase": "final_answer",

504 "content": "Root cause: cache invalidation race.",

505 },

506 {

507 "role": "user",

508 "content": "Great—now give me a rollout-safe fix plan.",

509 },

510 ],

511)

512 

513print(response.output_text)

514```

515 

516```go

517package main

518 

519import (

520 "context"

521 "fmt"

522 

523 "github.com/openai/openai-go/v3"

524 "github.com/openai/openai-go/v3/responses"

525)

526 

527func main() {

528 client := openai.NewClient()

529 commentary := responses.ResponseInputItemParamOfMessage(

530 "I’ll inspect the logs and then summarize root cause and remediation.",

531 responses.EasyInputMessageRoleAssistant,

532 )

533 commentary.OfMessage.Phase = responses.EasyInputMessagePhaseCommentary

534 finalAnswer := responses.ResponseInputItemParamOfMessage(

535 "Root cause: cache invalidation race.",

536 responses.EasyInputMessageRoleAssistant,

537 )

538 finalAnswer.OfMessage.Phase = responses.EasyInputMessagePhaseFinalAnswer

539 

540 response, err := client.Responses.New(context.Background(), responses.ResponseNewParams{

541 Model: "gpt-5.4",

542 Input: responses.ResponseNewParamsInputUnion{OfInputItemList: responses.ResponseInputParam{

543 commentary,

544 finalAnswer,

545 responses.ResponseInputItemParamOfMessage("Great—now give me a rollout-safe fix plan.", responses.EasyInputMessageRoleUser),

546 }},

547 })

548 if err != nil {

549 panic(err)

550 }

551 fmt.Println(response.OutputText())

552}

553```

554 

555```java

556import com.openai.client.OpenAIClient;

557import com.openai.client.okhttp.OpenAIOkHttpClient;

558import com.openai.models.Reasoning;

559import com.openai.models.ReasoningEffort;

560import com.openai.models.responses.ResponseCreateParams;

561 

562ResponseCreateParams params =

563 ResponseCreateParams.builder()

564 .model("gpt-5.4")

565 .input("Explain the bug and propose a fix.")

566 .reasoning(Reasoning.builder().effort(ReasoningEffort.MEDIUM).build())

567 .build();

568 

569client.responses().create(params).output().stream()

570 .flatMap(item -> item.message().stream())

571 .flatMap(message -> message.content().stream())

572 .flatMap(content -> content.outputText().stream())

573 .forEach(text -> System.out.println(text.text()));

574```

575 

576```ruby

577require "openai"

578 

579client = OpenAI::Client.new

580response = client.responses.create(

581 model: "gpt-5.4",

582 reasoning: { effort: :medium },

583 input: "Explain the bug and propose a fix."

584)

585puts(response.output_text)

586```

587 

588 

589### GPT-5.4 parameter compatibility

590 

591The following parameters are **only supported** when using GPT-5.4 with reasoning effort set to `none`:

592 

593- `temperature`

594- `top_p`

595- `logprobs`

596 

597Requests that include these fields will raise an error for GPT-5.4 or GPT-5.2 with any other reasoning effort setting, or for older GPT-5 models such as `gpt-5`, `gpt-5-mini`, or `gpt-5-nano`.

598 

599To achieve similar results with reasoning effort set higher, or with another GPT-5 family model, try these alternative parameters:

600 

601- **Reasoning depth:** `reasoning: { effort: "none" | "low" | "medium" | "high" | "xhigh" }`

602- **Output verbosity:** `text: { verbosity: "low" | "medium" | "high" }`

603- **Output length:** `max_output_tokens`

604 

605### Migrating from Chat Completions to Responses API

606 

607The biggest difference, and main reason to migrate from Chat Completions to the Responses API for GPT-5.4, is support for passing chain of thought (CoT) between turns. See a full [comparison of the APIs](https://developers.openai.com/api/docs/guides/migrate-to-responses).

608 

609Passing CoT exists only in the Responses API, and we've seen improved intelligence, fewer generated reasoning tokens, higher cache hit rates, and lower latency as a result of doing so. Most other parameters remain at parity, though the formatting is different. Here's how new parameters are handled differently between Chat Completions and the Responses API:

610 

611**Reasoning effort**

612 

613 

614 

615Responses API

616 

617 Generate response with reasoning effort set to none

618 

619```bash

620curl --request POST \

621 --url https://api.openai.com/v1/responses \

622 --header "Authorization: Bearer $OPENAI_API_KEY" \

623 --header "Content-type: application/json" \

624 --data '{

625 "model": "gpt-5.4",

626 "input": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?",

627 "reasoning": {

628 "effort": "none"

629 }

630}'

631```

632 

633

634 

635

636 

637

638Chat Completions

639 

640 Generate response with reasoning effort set to none

641 

642```bash

643curl --request POST \

644 --url https://api.openai.com/v1/chat/completions \

645 --header "Authorization: Bearer $OPENAI_API_KEY" \

646 --header "Content-type: application/json" \

647 --data '{

648 "model": "gpt-5.4",

649 "messages": [

650 {

651 "role": "user",

652 "content": "How much gold would it take to coat the Statue of Liberty in a 1mm layer?"

653 }

654 ],

655 "reasoning_effort": "none"

656}'

657```

658 

659 

660 

661**Verbosity**

662 

663 

664 

665Responses API

666 

667 Control verbosity

668 

669```bash

670curl --request POST \

671 --url https://api.openai.com/v1/responses \

672 --header "Authorization: Bearer $OPENAI_API_KEY" \

673 --header "Content-type: application/json" \

674 --data '{

675 "model": "gpt-5.4",

676 "input": "What is the answer to the ultimate question of life, the universe, and everything?",

677 "text": {

678 "verbosity": "low"

679 }

680}'

681```

682 

683

684 

685

686 

687

688Chat Completions

689 

690 Control verbosity

691 

692```bash

693curl --request POST \

694 --url https://api.openai.com/v1/chat/completions \

695 --header "Authorization: Bearer $OPENAI_API_KEY" \

696 --header "Content-type: application/json" \

697 --data '{

698 "model": "gpt-5.4",

699 "messages": [

700 {

701 "role": "user",

702 "content": "What is the answer to the ultimate question of life, the universe, and everything?"

703 }

704 ],

705 "verbosity": "low"

706}'

707```

708 

709 

710 

711**Custom tools**

712 

713 

714 

715Responses API

716 

717 Custom tool call

718 

719```bash

720curl --request POST \

721 --url https://api.openai.com/v1/responses \

722 --header "Authorization: Bearer $OPENAI_API_KEY" \

723 --header "Content-type: application/json" \

724 --data '{

725 "model": "gpt-5.4",

726 "input": "Use the code_exec tool to calculate the area of a circle with radius equal to the number of r letters in blueberry",

727 "tools": [

728 {

729 "type": "custom",

730 "name": "code_exec",

731 "description": "Executes arbitrary Python code"

732 }

733 ]

734}'

735```

736 

737

738 

739

740 

741

742Chat Completions

743 

744 Custom tool call

745 

746```bash

747curl --request POST \

748 --url https://api.openai.com/v1/chat/completions \

749 --header "Authorization: Bearer $OPENAI_API_KEY" \

750 --header "Content-type: application/json" \

751 --data '{

752 "model": "gpt-5.4",

753 "messages": [

754 {

755 "role": "user",

756 "content": "Use the code_exec tool to calculate the area of a circle with radius equal to the number of r letters in blueberry"

757 }

758 ],

759 "tools": [

760 {

761 "type": "custom",

762 "custom": {

763 "name": "code_exec",

764 "description": "Executes arbitrary Python code"

765 }

766 }

767 ]

768}'

769```

770 

771 

772 

773 

774## Prompting best practices

775 

776When troubleshooting cases where GPT-5.4 treats an intermediate update as the

777 final answer, verify your integration preserves the assistant message `phase`

778 field correctly. See [Phase parameter](#phase-parameter) for details.

779 

780### Understand GPT-5.4 behavior

781 

782#### Where GPT-5.4 is strongest

783 

784GPT-5.4 tends to work especially well in these areas:

785 

786- Strong personality and tone adherence, with less drift over long answers

787- Agentic workflow robustness, with a stronger tendency to stick with multi-step work, retry, and complete agent loops end to end

788- Evidence-rich synthesis, especially in long-context or multi-tool workflows

789- Instruction adherence in modular, skill-based, and block-structured prompts when the contract is explicit

790- Long-context analysis across large, messy, or multi-document inputs

791- Batched or parallel tool calling while maintaining tool-call accuracy

792- Spreadsheet, finance, and Excel workflows that need instruction following, formatting fidelity, and stronger self-verification

793 

794#### Where explicit prompting still helps

795 

796Even with those strengths, GPT-5.4 benefits from more explicit guidance in a few recurring patterns:

797 

798- Low-context tool routing early in a session, when tool selection can be less reliable

799- Dependency-aware workflows that need explicit prerequisite and downstream-step checks

800- Reasoning effort selection, where higher effort is not always better and the right choice depends on task shape, not intuition

801- Research tasks that require disciplined source collection and consistent citations

802- Irreversible or high-impact actions that require verification before execution

803- Terminal or coding-agent environments where tool boundaries must stay clear

804 

805These patterns are observed defaults, not guarantees. Start with the smallest prompt that passes your evals, and add blocks only when they fix a measured failure mode.

806 

807### Use core prompt patterns

808 

809#### Keep outputs compact and structured

810 

811To improve token efficiency with GPT-5.4, constrain verbosity and enforce structured output through clear output contracts. In practice, this acts as an additional control layer alongside the `verbosity` parameter in the Responses API, allowing you to guide both how much the model writes and how it structures the output.

812 

813```xml

814<output_contract>

815- Return exactly the sections requested, in the requested order.

816- If the prompt defines a preamble, analysis block, or working section, do not treat it as extra output.

817- Apply length limits only to the section they are intended for.

818- If a format is required (JSON, Markdown, SQL, XML), output only that format.

819</output_contract>

820 

821<verbosity_controls>

822- Prefer concise, information-dense writing.

823- Avoid repeating the user's request.

824- Keep progress updates brief.

825- Do not shorten the answer so aggressively that required evidence, reasoning, or completion checks are omitted.

826</verbosity_controls>

827```

828 

829#### Set clear defaults for follow-through

830 

831Users often change the task, format, or tone mid-conversation. To keep the assistant aligned, define clear rules for when to proceed, when to ask, and how newer instructions override earlier defaults.

832 

833Use a default follow-through policy like this:

834 

835```xml

836<default_follow_through_policy>

837- If the user’s intent is clear and the next step is reversible and low-risk, proceed without asking.

838- Ask permission only if the next step is:

839 (a) irreversible,

840 (b) has external side effects (for example sending, purchasing, deleting, or writing to production), or

841 (c) requires missing sensitive information or a choice that would materially change the outcome.

842- If proceeding, briefly state what you did and what remains optional.

843</default_follow_through_policy>

844```

845 

846Make instruction priority explicit:

847 

848```xml

849<instruction_priority>

850- User instructions override default style, tone, formatting, and initiative preferences.

851- Safety, honesty, privacy, and permission constraints do not yield.

852- If a newer user instruction conflicts with an earlier one, follow the newer instruction.

853- Preserve earlier instructions that do not conflict.

854</instruction_priority>

855```

856 

857Higher-priority developer or system instructions remain binding.

858 

859**Guidance:** When instructions change mid-conversation, make the update explicit, scoped, and local. State what changed, what still applies, and whether the change affects the next turn or the rest of the conversation.

860 

861#### Handle mid-conversation instruction updates

862 

863For mid-conversation updates, use explicit, scoped steering messages that state:

864 

8651. Scope

8662. Override

8673. Carry forward

868 

869```text

870<task_update>

871For the next response only:

872- Do not complete the task.

873- Only produce a plan.

874- Keep it to 5 bullets.

875 

876All earlier instructions still apply unless they conflict with this update.

877</task_update>

878```

879 

880If the task itself changes, say so directly:

881 

882```text

883<task_update>

884The task has changed.

885Previous task: complete the workflow.

886Current task: review the workflow and identify risks only.

887 

888Rules for this turn:

889- Do not execute actions.

890- Do not call destructive tools.

891- Return exactly:

892 1. Main risks

893 2. Missing information

894 3. Recommended next step

895</task_update>

896```

897 

898#### Make tool use persistent when correctness depends on it

899 

900Use explicit rules to keep tool use thorough, dependency-aware, and appropriately paced, especially in workflows where later actions rely on earlier retrieval or verification. A common failure mode is skipping prerequisites because the right end state seems obvious.

901 

902GPT-5.4 can be less reliable at tool routing early in a session, when context is still thin. Prompt for prerequisites, dependency checks, and exact tool intent.

903 

904```xml

905<tool_persistence_rules>

906- Use tools whenever they materially improve correctness, completeness, or grounding.

907- Do not stop early when another tool call is likely to materially improve correctness or completeness.

908- Keep calling tools until:

909 (1) the task is complete, and

910 (2) verification passes (see <verification_loop>).

911- If a tool returns empty or partial results, retry with a different strategy.

912</tool_persistence_rules>

913```

914 

915This is especially important for workflows where the final action depends on earlier lookup or retrieval steps. One of the most common failure modes is skipping prerequisites because the intended end state seems obvious.

916 

917```xml

918<dependency_checks>

919- Before taking an action, check whether prerequisite discovery, lookup, or memory retrieval steps are required.

920- Do not skip prerequisite steps just because the intended final action seems obvious.

921- If the task depends on the output of a prior step, resolve that dependency first.

922</dependency_checks>

923```

924 

925Prompt for parallelism when the work is independent and wall-clock matters. Prompt for sequencing when dependencies, ambiguity, or irreversible actions matter more than speed.

926 

927```xml

928<parallel_tool_calling>

929- When multiple retrieval or lookup steps are independent, prefer parallel tool calls to reduce wall-clock time.

930- Do not parallelize steps that have prerequisite dependencies or where one result determines the next action.

931- After parallel retrieval, pause to synthesize the results before making more calls.

932- Prefer selective parallelism: parallelize independent evidence gathering, not speculative or redundant tool use.

933</parallel_tool_calling>

934```

935 

936#### Force completeness on long-horizon tasks

937 

938For multi-step workflows, a common failure mode is incomplete execution: the model finishes after partial coverage, misses items in a batch, or treats empty or narrow retrieval as final. GPT-5.4 becomes more reliable when the prompt defines explicit completion rules and recovery behavior.

939 

940Coverage can be achieved through sequential or parallel retrieval, but completion rules should remain explicit either way.

941 

942```xml

943<completeness_contract>

944- Treat the task as incomplete until all requested items are covered or explicitly marked [blocked].

945- Keep an internal checklist of required deliverables.

946- For lists, batches, or paginated results:

947 - determine expected scope when possible,

948 - track processed items or pages,

949 - confirm coverage before finalizing.

950- If any item is blocked by missing data, mark it [blocked] and state exactly what is missing.

951</completeness_contract>

952```

953 

954For workflows where empty, partial, or noisy retrieval is common:

955 

956```xml

957<empty_result_recovery>

958If a lookup returns empty, partial, or suspiciously narrow results:

959- do not immediately conclude that no results exist,

960- try at least one or two fallback strategies,

961 such as:

962 - alternate query wording,

963 - broader filters,

964 - a prerequisite lookup,

965 - or an alternate source or tool,

966- Only then report that no results were found, along with what you tried.

967</empty_result_recovery>

968```

969 

970#### Add a verification loop before high-impact actions

971 

972Once the workflow appears complete, add a lightweight verification step before returning the answer or taking an irreversible action. This helps catch requirement misses, grounding issues, and format drift before commit.

973 

974```xml

975<verification_loop>

976Before finalizing:

977- Check correctness: does the output satisfy every requirement?

978- Check grounding: are factual claims backed by the provided context or tool outputs?

979- Check formatting: does the output match the requested schema or style?

980- Check safety and irreversibility: if the next step has external side effects, ask permission first.

981</verification_loop>

982```

983 

984```xml

985<missing_context_gating>

986- If required context is missing, do NOT guess.

987- Prefer the appropriate lookup tool when the missing context is retrievable; ask a minimal clarifying question only when it is not.

988- If you must proceed, label assumptions explicitly and choose a reversible action.

989</missing_context_gating>

990```

991 

992For agents that actively take actions, add a short execution frame:

993 

994```xml

995<action_safety>

996- Pre-flight: summarize the intended action and parameters in 1-2 lines.

997- Execute via tool.

998- Post-flight: confirm the outcome and any validation that was performed.

999</action_safety>

1000```

1001 

1002### Handle specialized workflows

1003 

1004#### Choose image detail explicitly for vision and computer use

1005 

1006If your workflow depends on visual precision, specify the image `detail` level in the prompt or integration instead of relying on `auto`. Use `high` for standard high-fidelity image understanding. Use `original` for large, dense, or spatially sensitive images, especially [computer use, localization, OCR, and click-accuracy tasks](https://developers.openai.com/api/docs/guides/tools-computer-use) on `gpt-5.4` and future models. Use `low` only when speed and cost matter more than fine detail. For more details on image detail levels, see the [Images and Vision guide](https://developers.openai.com/api/docs/guides/images-vision).

1007 

1008#### Lock research and citations to retrieved evidence

1009 

1010When citation quality matters, make both the source boundary and the format requirement explicit. This helps reduce fabricated references, unsupported claims, and citation-format drift.

1011 

1012```xml

1013<citation_rules>

1014- Only cite sources retrieved in the current workflow.

1015- Never fabricate citations, URLs, IDs, or quote spans.

1016- Use exactly the citation format required by the host application.

1017- Attach citations to the specific claims they support, not only at the end.

1018</citation_rules>

1019```

1020 

1021```xml

1022<grounding_rules>

1023- Base claims only on provided context or tool outputs.

1024- If sources conflict, state the conflict explicitly and attribute each side.

1025- If the context is insufficient or irrelevant, narrow the answer or say you cannot support the claim.

1026- If a statement is an inference rather than a directly supported fact, label it as an inference.

1027</grounding_rules>

1028```

1029 

1030If your application requires inline citations, require inline citations. If it requires footnotes, require footnotes. The key is to lock the format and prevent the model from improvising unsupported references.

1031 

1032#### Research mode

1033 

1034Push GPT-5.4 into a disciplined research mode. Use this pattern for research, review, and synthesis tasks. Do not force it onto short execution tasks or simple deterministic transforms.

1035 

1036```xml

1037<research_mode>

1038- Do research in 3 passes:

1039 1) Plan: list 3-6 sub-questions to answer.

1040 2) Retrieve: search each sub-question and follow 1-2 second-order leads.

1041 3) Synthesize: resolve contradictions and write the final answer with citations.

1042- Stop only when more searching is unlikely to change the conclusion.

1043</research_mode>

1044```

1045 

1046If your host environment uses a specific research tool or requires a submit step, combine this with the host's finalization contract.

1047 

1048#### Clamp strict output formats

1049 

1050For SQL, JSON, or other parse-sensitive outputs, tell GPT-5.4 to emit only the target format and check it before finishing.

1051 

1052```text

1053<structured_output_contract>

1054- Output only the requested format.

1055- Do not add prose or markdown fences unless they were requested.

1056- Validate that parentheses and brackets are balanced.

1057- Do not invent tables or fields.

1058- If required schema information is missing, ask for it or return an explicit error object.

1059</structured_output_contract>

1060```

1061 

1062If you are extracting document regions or OCR boxes, define the coordinate system and add a drift check:

1063 

1064```text

1065<bbox_extraction_spec>

1066- Use the specified coordinate format exactly, such as [x1,y1,x2,y2] normalized to 0..1.

1067- For each box, include page, label, text snippet, and confidence.

1068- Add a vertical-drift sanity check so boxes stay aligned with the correct line of text.

1069- If the layout is dense, process page by page and do a second pass for missed items.

1070</bbox_extraction_spec>

1071```

1072 

1073#### Keep tool boundaries explicit in coding and terminal agents

1074 

1075In coding agents, GPT-5.4 works better when the rules for shell access and file editing are unambiguous. This is especially important when you expose tools like [Shell](https://developers.openai.com/api/docs/guides/tools-shell) or [Apply patch](https://developers.openai.com/api/docs/guides/tools-apply-patch).

1076 

1077#### User updates

1078 

1079GPT-5.4 does well with brief, outcome-based updates. Reuse the user-updates pattern from the 5.2 guide, but pair it with explicit completion and verification requirements.

1080 

1081Recommended update spec:

1082 

1083```xml

1084<user_updates_spec>

1085- Only update the user when starting a new major phase or when something changes the plan.

1086- Each update: 1 sentence on outcome + 1 sentence on next step.

1087- Do not narrate routine tool calls.

1088- Keep the user-facing status short; keep the work exhaustive.

1089</user_updates_spec>

1090```

1091 

1092For coding agents, see the Prompting patterns for coding tasks section below for more specific guidance.

1093 

1094#### Prompting patterns for coding tasks

1095 

1096**Autonomy and persistence**

1097 

1098GPT-5.4 is generally more thorough end to end than earlier mainline models on coding and tool-use tasks, so you often need less explicit "verify everything" prompting. Still, for high-stakes changes such as production, migrations, or security work, keep a lightweight verification clause.

1099 

1100```xml

1101<autonomy_and_persistence>

1102Persist until the task is fully handled end-to-end within the current turn whenever feasible: do not stop at analysis or partial fixes; carry changes through implementation, verification, and a clear explanation of outcomes unless the user explicitly pauses or redirects you.

1103 

1104Unless the user explicitly asks for a plan, asks a question about the code, is brainstorming potential solutions, or some other intent that makes it clear that code should not be written, assume the user wants you to make code changes or run tools to solve the user's problem. In these cases, it's bad to output your proposed solution in a message, you should go ahead and actually implement the change. If you encounter challenges or blockers, you should attempt to resolve them yourself.

1105</autonomy_and_persistence>

1106```

1107 

1108**Intermediary updates**

1109 

1110Keep updates sparse and high-signal. In coding tasks, prefer updates at key points.

1111 

1112```xml

1113<user_updates_spec>

1114- Intermediary updates go to the `commentary` channel.

1115- User updates are short updates while you are working. They are not final answers.

1116- Use 1-2 sentence updates to communicate progress and new information while you work.

1117- Do not begin responses with conversational interjections or meta commentary. Avoid openers such as acknowledgements ("Done -", "Got it", or "Great question") or similar framing.

1118- Before exploring or doing substantial work, send a user update explaining your understanding of the request and your first step. Avoid commenting on the request or starting with phrases such as "Got it" or "Understood."

1119- Provide updates roughly every 30 seconds while working.

1120- When exploring, explain what context you are gathering and what you learned. Vary sentence structure so the updates do not become repetitive.

1121- When working for a while, keep updates informative and varied, but stay concise.

1122- When work is substantial, provide a longer plan after you have enough context. This is the only update that may be longer than 2 sentences and may contain formatting.

1123- Before file edits, explain what you are about to change.

1124- While thinking, keep the user informed of progress without narrating every tool call. Even if you are not taking actions, send frequent progress updates rather than going silent, especially if you are thinking for more than a short stretch.

1125- Keep the tone of progress updates consistent with the assistant's overall personality.

1126</user_updates_spec>

1127```

1128 

1129**Formatting**

1130 

1131GPT-5.4 often defaults to more structured formatting and may overuse bullet lists. If you want a clean final response, explicitly clamp list shape.

1132 

1133```xml

1134Never use nested bullets. Keep lists flat (single level). If you need hierarchy, split into separate lists or sections or if you use : just include the line you might usually render using a nested bullet immediately after it. For numbered lists, only use the `1. 2. 3.` style markers (with a period), never `1)`.

1135```

1136 

1137**Frontend tasks**

1138 

1139Use this only when additional frontend guidance is useful.

1140 

1141```xml

1142<frontend_tasks>

1143When doing frontend design tasks, avoid generic, overbuilt layouts.

1144 

1145Use these hard rules:

1146- One composition: The first viewport must read as one composition, not a dashboard, unless it is a dashboard.

1147- Brand first: On branded pages, the brand or product name must be a hero-level signal, not just nav text or an eyebrow. No headline should overpower the brand.

1148- Brand test: If the first viewport could belong to another brand after removing the nav, the branding is too weak.

1149- Full-bleed hero only: On landing pages and promotional surfaces, the hero image should usually be a dominant edge-to-edge visual plane or background. Do not default to inset hero images, side-panel hero images, rounded media cards, tiled collages, or floating image blocks unless the existing design system clearly requires them.

1150- Hero budget: The first viewport should usually contain only the brand, one headline, one short supporting sentence, one CTA group, and one dominant image. Do not place stats, schedules, event listings, address blocks, promos, "this week" callouts, metadata rows, or secondary marketing content there.

1151- No hero overlays: Do not place detached labels, floating badges, promo stickers, info chips, or callout boxes on top of hero media.

1152- Cards: Default to no cards. Never use cards in the hero unless they are the container for a user interaction. If removing a border, shadow, background, or radius does not hurt interaction or understanding, it should not be a card.

1153- One job per section: Each section should have one purpose, one headline, and usually one short supporting sentence.

1154- Real visual anchor: Imagery should show the product, place, atmosphere, or context.

1155- Reduce clutter: Avoid pill clusters, stat strips, icon rows, boxed promos, schedule snippets, and competing text blocks.

1156- Use motion to create presence and hierarchy, not noise. Ship 2-3 intentional motions for visually led work, and prefer Framer Motion when it is available.

1157 

1158Exception: If working within an existing website or design system, preserve the established patterns, structure, and visual language.

1159</frontend_tasks>

1160```

1161 

1162```xml

1163<terminal_tool_hygiene>

1164- Only run shell commands via the terminal tool.

1165- Never "run" tool names as shell commands.

1166- If a patch or edit tool exists, use it directly; do not attempt it in bash.

1167- After changes, run a lightweight verification step such as ls, tests, or a build before declaring the task done.

1168</terminal_tool_hygiene>

1169```

1170 

1171#### Document localization and OCR boxes

1172 

1173For bbox tasks, be explicit about coordinate conventions and add drift tests.

1174 

1175```xml

1176<bbox_extraction_spec>

1177- Use the specified coordinate format exactly (for example [x1,y1,x2,y2] normalized 0..1).

1178- For each bbox, include: page, label, text snippet, confidence.

1179- Add a vertical-drift sanity check:

1180 - ensure bboxes align with the line of text (not shifted up or down).

1181- If dense layout, process page by page and do a second pass for missed items.

1182</bbox_extraction_spec>

1183```

1184 

1185#### Use runtime and API integration notes

1186 

1187For long-running or tool-heavy agents, the runtime contract matters as much as the prompt contract.

1188 

1189##### Phase parameter

1190 

1191For GPT-5.4, `gpt-5.3-codex`, and later Responses models, the `phase` field can

1192help in the small number of long-running or tool-heavy flows where preambles or

1193other intermediate assistant updates are mistaken for the final answer.

1194 

1195- `phase` is optional at the API level, but it is highly recommended. Best-effort inference may exist server-side, but explicit round-tripping of `phase` is strictly better.

1196- Use `phase` for long-running or tool-heavy agents that may emit commentary before tool calls or before a final answer.

1197- Preserve `phase` when replaying prior assistant items so the model can distinguish working commentary from the completed answer. This matters most in multi-step flows with preambles, tool-related updates, or multiple assistant messages in the same turn.

1198- Do not add `phase` to user messages.

1199- If you use `previous_response_id`, that is usually the simplest path, since OpenAI can often recover prior state without manually replaying assistant items.

1200- If you replay assistant history yourself, preserve the original `phase` values.

1201- Missing or dropped `phase` can cause preambles to be interpreted as final answers and degrade behavior on those multi-step tasks.

1202 

1203#### Preserve behavior in long sessions

1204 

1205Compaction unlocks significantly longer effective context windows, where user conversations can persist for many turns without hitting context limits or long-context performance degradation, and agents can perform very long trajectories that exceed a typical context window for long-running, complex tasks.

1206 

1207If you are using [Compaction](https://developers.openai.com/api/docs/guides/compaction) in the Responses API, compact after major milestones, treat compacted items as opaque state, and keep prompts functionally identical after compaction. The endpoint is ZDR compatible and returns an `encrypted_content` item that you can pass into future requests. GPT-5.4 tends to remain more coherent and reliable over longer, multi-turn conversations with fewer breakdowns as sessions grow.

1208 

1209For more guidance, see the [`/responses/compact` API reference](https://developers.openai.com/api/reference/resources/responses/methods/compact).

1210 

1211#### Control personality for customer-facing workflows

1212 

1213GPT-5.4 can be steered more effectively when you separate persistent personality from per-response writing controls. This is especially useful for customer-facing workflows such as emails, support replies, announcements, and blog-style content.

1214 

1215- **Personality (persistent):** sets the default tone, verbosity, and decision style across the session.

1216- **Writing controls (per response):** define the channel, register, formatting, and length for a specific artifact.

1217- **Reminder:** personality should not override task-specific output requirements. If the user asks for JSON, return JSON.

1218 

1219For natural, high-quality prose, the highest-leverage controls are:

1220 

1221- Give the model a clear persona.

1222- Specify the channel and emotional register.

1223- Explicitly ban formatting when you want prose.

1224- Use hard length limits.

1225 

1226```xml

1227<personality_and_writing_controls>

1228- Persona: <one sentence>

1229- Channel: <Slack | email | memo | PRD | blog>

1230- Emotional register: <direct/calm/energized/etc.> + "not <overdo this>"

1231- Formatting: <ban bullets/headers/markdown if you want prose>

1232- Length: <hard limit, e.g. <=150 words or 3-5 sentences>

1233- Default follow-through: if the request is clear and low-risk, proceed without asking permission.

1234</personality_and_writing_controls>

1235```

1236 

1237For more personality patterns you can lift directly, see the [Prompt Personalities cookbook](https://developers.openai.com/cookbook/examples/gpt-5/prompt_personalities).

1238 

1239**Professional memo mode**

1240 

1241For memos, reviews, and other professional writing tasks, general writing instructions are often not enough. These workflows benefit from explicit guidance on specificity, domain conventions, synthesis, and calibrated certainty.

1242 

1243```xml

1244<memo_mode>

1245- Write in a polished, professional memo style.

1246- Use exact names, dates, entities, and authorities when supported by the record.

1247- Follow domain-specific structure if one is requested.

1248- Prefer precise conclusions over generic hedging.

1249- When uncertainty is real, tie it to the exact missing fact or conflicting source.

1250- Synthesize across documents rather than summarizing each one independently.

1251</memo_mode>

1252```

1253 

1254This mode is especially useful for legal, policy, research, and executive-facing writing, where the goal is not just fluency, but disciplined synthesis and clear conclusions.

1255 

1256### Tune reasoning and migration

1257 

1258#### Treat reasoning effort as a last-mile knob

1259 

1260Reasoning effort is not one-size-fits-all. Treat it as a last-mile tuning knob, not the primary way to improve quality. In many cases, stronger prompts, clear output contracts, and lightweight verification loops recover much of the performance teams might otherwise seek through higher reasoning settings.

1261 

1262Recommended defaults:

1263 

1264- `none`: Best for fast, cost-sensitive, latency-sensitive tasks where the model does not need to think.

1265- `low`: Works well for latency-sensitive tasks where a small amount of thinking can produce a meaningful accuracy gain, especially with complex instructions.

1266- `medium` or `high`: Reserve for tasks that truly require stronger reasoning and can absorb the latency and cost tradeoff. Choose between them based on how much performance gain your task gets from additional reasoning.

1267- `xhigh`: Avoid as a default unless your evals show clear benefits. It is best suited for long, agentic, reasoning-heavy tasks where maximum intelligence matters more than speed or cost.

1268 

1269In practice, most teams should default to the `none`, `low`, or `medium` range.

1270 

1271Start with `none` for execution-heavy workloads such as workflow steps, field extraction, support triage, and short structured transforms.

1272 

1273Start with `medium` or higher for research-heavy workloads such as long-context synthesis, multi-document review, conflict resolution, and strategy writing. With `medium` and a well-engineered prompt, you can squeeze out a lot of performance.

1274 

1275For GPT-5.4 workloads, `none` can already perform well on action-selection and tool-discipline tasks. If your workload depends on nuanced interpretation, such as implicit requirements, ambiguity, or cancelled-tool-call recovery, start with `low` or `medium` instead.

1276 

1277Before increasing reasoning effort, first add:

1278 

1279- `<completeness_contract>`

1280- `<verification_loop>`

1281- `<tool_persistence_rules>`

1282 

1283If the model still feels too literal or stops at the first plausible answer, add an initiative nudge before raising reasoning effort:

1284 

1285```xml

1286<dig_deeper_nudge>

1287- Don’t stop at the first plausible answer.

1288- Look for second-order issues, edge cases, and missing constraints.

1289- If the task is safety or accuracy critical, perform at least one verification step.

1290</dig_deeper_nudge>

1291```

1292 

1293#### Migrate prompts to GPT-5.4 one change at a time

1294 

1295Use the same one-change-at-a-time discipline as the 5.2 guide: switch model first, pin `reasoning_effort`, run evals, then iterate.

1296 

1297These starting points work well for many migrations:

1298 

1299| Current setup | Suggested GPT-5.4 start | Notes |

1300| ------------------------- | ---------------------------------- | ------------------------------------------------------------------- |

1301| `gpt-5.2` | Match the current reasoning effort | Preserve the existing latency and quality profile first, then tune. |

1302| `gpt-5.3-codex` | Match the current reasoning effort | For coding workflows, keep the reasoning effort the same. |

1303| `gpt-4.1` or `gpt-4o` | `none` | Keep snappy behavior, and increase only if evals regress. |

1304| Research-heavy assistants | `medium` or `high` | Use explicit research multi-pass and citation gating. |

1305| Long-horizon agents | `medium` or `high` | Add tool persistence and completeness accounting. |

1306 

1307#### Small-model guidance for `gpt-5.4-mini` and `gpt-5.4-nano`

1308 

1309`gpt-5.4-mini` and `gpt-5.4-nano` are highly steerable, but they are less likely than larger models to infer missing steps, resolve ambiguity implicitly, or package outputs the way you intended unless you specify that behavior directly. In practice, prompts for smaller models are often a bit longer and more explicit.

1310 

1311**How `gpt-5.4-mini` differs**

1312 

1313- `gpt-5.4-mini` is more literal and makes fewer assumptions.

1314- It is strong when the task is clearly structured, but weaker on implicit workflows and ambiguity handling.

1315- By default, it may try to keep the conversation going with a follow-up question unless you suppress that behavior explicitly.

1316 

1317**Prompting `gpt-5.4-mini`**

1318 

1319- Put critical rules first.

1320- Specify the full execution order when tool use or side effects matter.

1321- Do not rely on "you MUST" alone. Use structural scaffolding such as numbered steps, decision rules, and explicit action definitions.

1322- Separate "do the action" from "report the action."

1323- Show the correct flow, not just the final format.

1324- Define ambiguity behavior explicitly: when to ask, abstain, or proceed.

1325- Specify packaging directly: answer length, whether to ask a follow-up question, citation style, and section order.

1326- Be careful with `output nothing else`. Prefer scoped instructions such as `after the final JSON, output nothing further`.

1327 

1328**Prompting `gpt-5.4-nano`**

1329 

1330- Use `gpt-5.4-nano` only for narrow, well-bounded tasks.

1331- Prefer closed outputs: labels, enums, short JSON, or fixed templates.

1332- Avoid multi-step orchestration unless the flow is extremely constrained.

1333- Route ambiguous or planning-heavy tasks to a stronger model instead of over-prompting `gpt-5.4-nano`.

1334 

1335**Good default pattern**

1336 

13371. Task

13382. Critical rule

13393. Exact step order

13404. Edge cases or clarification behavior

13415. Output format

13426. One correct example

1343 

1344**Avoid**

1345 

1346- Implied next steps

1347- Unspecified edge cases

1348- Schema-only prompts for tool workflows

1349- Generic instructions without structure

1350 

1351#### Web search and deep research

1352 

1353If you are migrating a research agent in particular, make these prompt updates before increasing reasoning effort:

1354 

1355- Add `<research_mode>`

1356- Add `<citation_rules>`

1357- Add `<empty_result_recovery>`

1358- Increase `reasoning_effort` one notch only after prompt fixes.

1359 

1360You can start from the 5.2 research block and then layer in citation gating and finalization contracts as needed.

1361 

1362GPT-5.4 performs especially well when the task requires multi-step evidence gathering, long-context synthesis, and explicit prompt contracts. In practice, the highest-leverage prompt changes are choosing reasoning effort by task shape, defining exact output and citation formats, adding dependency-aware tool rules, and making completion criteria explicit. The model is often strong out of the box, but it is most reliable when prompts clearly specify how to search, how to verify, and what counts as done.

1363 

1364### Next steps

1365 

1366- Review [Model, API, and feature updates](#model-api-and-feature-updates) for model capabilities, parameters, and API compatibility details.

1367- Read [Prompt engineering](https://developers.openai.com/api/docs/guides/prompt-engineering) for broader prompting strategies that apply across model families.

1368- Read [Compaction](https://developers.openai.com/api/docs/guides/compaction) if you are building long-running GPT-5.4 sessions in the Responses API.

1369 

1370 

1371## Further reading

1372 

1373[GPT-5.3-Codex prompting guide](https://developers.openai.com/cookbook/examples/gpt-5/codex_prompting_guide)

1374 

1375[GPT-5.4 blog post](https://openai.com/index/introducing-gpt-5-4/)

1376 

1377[GPT-5 frontend guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_frontend)

1378 

1379[GPT-5 model family: new features guide](https://developers.openai.com/cookbook/examples/gpt-5/gpt-5_new_params_and_tools)

1380 

1381[Cookbook on reasoning models](https://developers.openai.com/cookbook/examples/responses_api/reasoning_items)

1382 

1383[Comparison of Responses API vs. Chat Completions](https://developers.openai.com/api/docs/guides/migrate-to-responses)

guides/latest-model/gpt-5.5.md +0 −328 deleted

File Deleted View Diff

1# Using GPT-5.5

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5## Introduction

6 

7GPT-5.5 raises the baseline for complex production workflows. It’s a strong fit for coding use cases, tool-heavy agents, grounded assistants, long-context retrieval, product-spec-to-plan workflows, and customer-facing workflows where execution quality and response polish are critical.

8 

9To get the most out of GPT-5.5, treat it as a new model family to tune for, not a drop-in replacement for `gpt-5.2` or `gpt-5.4`. Begin migration with a fresh baseline instead of carrying over every instruction from an older prompt stack. Start with the smallest prompt that preserves the product contract, then tune reasoning effort, verbosity, tool descriptions, and output format against representative examples.

10 

11GPT-5.5 supports all API features that were already available with GPT-5.4, including [prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching), [hosted tools](https://developers.openai.com/api/docs/guides/tools#available-tools), [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search), [compaction](https://developers.openai.com/api/docs/guides/compaction), and `phase` handling for manually replayed assistant items.

12 

13See [Prompting best practices](#prompting-best-practices) for examples of successful prompting patterns.

14 

15## What's new

16 

17- **More efficient reasoning:** GPT-5.5 reaches strong results with fewer reasoning tokens than prior models, even at the same reasoning effort. This is especially useful in complex, tool-heavy, or multi-step workflows where token savings compound.

18- **Stronger task execution with outcome-first prompts:** GPT-5.5 is better at working from a clear goal, preserving constraints, and turning product intent into concrete next steps. Describe the expected outcome, success criteria, allowed side effects, evidence rules, and output shape. Avoid step-by-step process guidance unless the exact path matters.

19- **Stronger and more precise tool use:** GPT-5.5 is especially useful on large tool surfaces, multi-step service workflows, and long-running agent tasks. It tends to be more precise in tool selection and argument use.

20- **Tone is often more polished, but can be more direct:** GPT-5.5 often produces warmer, more readable answers with less prompt scaffolding.

21 

22## Behavioral changes

23 

241. **Reasoning effort now defaults to `medium`:** GPT-5.5 defaults to `medium` reasoning effort. Treat `medium` as the recommended balanced starting point for quality, reliability, latency, and cost. For latency-sensitive workflows, evaluate `low` before `none` when tool use, planning, search, or multi-step decision making still matters. Reserve `none` for latency-critical tasks that don't need reasoning or multi-chained tool calls, such as lightweight voice turns, fast information retrieval, and classification. Increase to `high` or `xhigh` only when evals show a measurable quality gain that justifies the extra latency and cost. See the [Reasoning models documentation](https://developers.openai.com/api/docs/guides/reasoning) for more details on recommended settings.

25 

26 Higher reasoning effort isn't automatically better. If the task has conflicting instructions, weak stopping criteria, or open-ended tool access, higher effort can lead to overthinking, unnecessary searching, or output quality regressions. Increase effort only when evals show a measurable quality gain.

27 

282. **Image inputs preserve more visual detail by default:** GPT-5.5 updates the default handling for image inputs to preserve more visual detail and improve computer use performance. When `image_detail` is unset or set to `auto`, the model now uses `original` behavior, preserving images without resizing up to 10,240,000 pixels or a 6,000-pixel dimension limit. For `high`, specify the value directly; it preserves images without resizing up to 2,500,000 pixels or a 2,048-pixel dimension limit. `low` now focuses on context efficiency and resizes images above a 512-pixel dimension limit more aggressively than previous models. See the [Images and vision documentation](https://developers.openai.com/api/docs/guides/images-vision).

29 

303. **Improved instruction following:** GPT-5.5 interprets prompts in a literal and thorough manner, enabling specific, descriptive instructions when the product requires them. Define success criteria and stopping rules, especially for long-running, tool-heavy, or evidence-gathering workflows. See [Write outcome-first prompts](#outcome-first-prompts-and-stopping-conditions) and [Keep the right specificity](#formatting).

31 

324. **Default style is more concise and direct:** GPT-5.5 tends to be efficient, direct, and task-oriented by default. This is useful for many production workflows, but customer-facing or conversational experiences may need explicit personality, warmth, rationale, and formatting guidance. Use `text.verbosity` intentionally: `medium` is the default, and `low` is often a better starting point for concise responses. See [Prompting best practices](#prompting-best-practices).

33 

345. **Coding workflows need stronger orchestration:** GPT-5.5 is better suited to complex coding tasks that require planning, tool use, codebase navigation, verification, and multi-step execution. For coding agents, be explicit about reuse, subagent delegation, test expectations, acceptance criteria, and when to continue versus ask for help.

35 

36## Migration quickstart

37 

38### Automated migration with Codex

39 

40Codex can apply the recommended changes in this guide with the [OpenAI Docs skill](https://github.com/openai/skills/tree/main/skills/.curated/openai-docs).

41 

42```text

43$openai-docs migrate this project to gpt-5.5

44```

45 

46To use this skill in other coding agents, download it from the [OpenAI skills repository](https://github.com/openai/skills/tree/main/skills/.curated/openai-docs).

47 

48### API and model parameters

49 

50- Update the model slug to `gpt-5.5`.

51- Use the Responses API for any reasoning, tool-calling, or multi-turn use case.

52- Tune `reasoning.effort`. Use `low` for efficient reasoning, `medium` for a balanced point on the latency/performance curve, `high` for complex agentic tasks that require hard reasoning and where latency matters less, and `xhigh` for the hardest asynchronous agentic tasks or evals that test the bounds of model intelligence. See the [Reasoning models documentation](https://developers.openai.com/api/docs/guides/reasoning).

53- To configure for more concise responses, set `text.verbosity` to `low`. On GPT-5.5, this will result in proportionally more concise responses than `low` verbosity with GPT-5.4.

54- For tool-heavy or long-running workflows, verify that your application handles `phase`, preambles, and assistant-item replay correctly.

55- Benchmark against other models on accuracy, token consumption, and end-to-end latency.

56 

57### Prompting

58 

59- State the expected outcome and success criteria.

60- Reduce or remove detailed step-by-step process guidance. Let GPT-5.5 choose the path unless the product requires that path.

61- Remove output schema definitions from the prompt where possible. Use [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) instead.

62- Optimize your prompt for caching: [static parts first, dynamic parts last](https://developers.openai.com/api/docs/guides/prompt-caching).

63- Drop the current date. The model is already aware of the current UTC date.

64- Review and optimize your prompts with [Prompting best practices](#prompting-best-practices).

65 

66## Using reasoning models

67 

68This guidance applies to GPT-5 series models and is worth revisiting whenever teams move workloads onto reasoning models. GPT-5.5 carries forward many capabilities that first appeared in earlier models, but they're still worth reviewing if you are moving from an earlier GPT-5 model, GPT-4.1, or a reasoning model such as o3.

69 

70Teams can overlook these features because they sit partly in API configuration and orchestration rather than in the prompt itself. Used together, the Responses API, reasoning controls, verbosity, structured outputs, prompt caching, tool design, hosted tools, and state management help reasoning models deliver their best intelligence, reliability, latency, and cost profile.

71 

72- **Responses API:** GPT-5.5 works best in the [Responses API](https://developers.openai.com/api/docs/guides/migrate-to-responses). Use `previous_response_id` for multi-turn state handling. For stateless or Zero Data Retention flows, pass back the relevant returned output items each turn. See [Passing context from the previous response](https://developers.openai.com/api/docs/guides/conversation-state#passing-context-from-the-previous-response) for details.

73- **Reasoning effort:** Use `reasoning.effort` to choose between `low`, `medium`, `high`, or `xhigh`. The default is `medium`, but many workloads will perform well with `low`. Reserve `none` for use cases where low latency is more important than intelligence. See [Reasoning Models](https://developers.openai.com/api/docs/guides/reasoning) for detailed recommendations.

74- **Verbosity:** Use `text.verbosity` to control output length. Treat final answer length as separate from reasoning quality; specify word budgets, section counts, table widths, or JSON-only output where needed.

75- **Structured Outputs:** Avoid describing the expected output schema in the prompt. Use [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) for automatic validation and increased accuracy.

76- **Prompt caching:** [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) works automatically for eligible long prompts and can reduce latency and input-token cost. To maximize cache hits, keep stable content at the beginning of the request. Put dynamic user-specific context near the end. Track `usage.prompt_tokens_details.cached_tokens` to measure reuse. Use a stable [`prompt_cache_key`](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-keys) for requests that share a reusable prefix. The key helps route related requests to the same cache and is important for optimizing cache hit rates on GPT-5.5. For busy groups, follow the [guidance for distributing traffic across more keys](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-keys).

77- **Tool calling:** GPT-5.5 supports the same tool-calling patterns as GPT-5.4, including function tools and tool-heavy agent workflows. Put most tool-specific guidance in the tool descriptions themselves: what the tool does, when to use it, required inputs, side effects, retry safety, and common error modes. Add tool-specific context to system instructions only when it applies across tools or materially changes the agent's operating policy.

78- **Hosted tools and tool search:** Prefer [OpenAI-hosted tools](https://developers.openai.com/api/docs/guides/tools) where they fit the workflow, such as web search, file search, code interpreter, image generation, and computer use. Hosted tools reduce custom orchestration burden and keep common tool patterns aligned with the Responses API and Agents SDK. Use custom function tools when you need to call your own systems, enforce domain-specific side effects, or expose internal business workflows. For large tool catalogs, consider using [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search) to defer tool definitions and load only the relevant subset.

79- **Tool preambles:** Preambles can improve chat UX because the user sees an initial, useful status update before the model generates the final response. They also make tool use easier to follow: the model can state what it's about to check or do, then continue from that same assistant state after tool results arrive.

80- **`phase` handling:** If your application manually manages Responses state by passing output items back each turn instead of using `previous_response_id`, preserve the `phase` parameter on returned assistant output items and pass it back unchanged. This is especially important when using reasoning effort, preambles, or repeated tool calls. See [Phase parameter](https://developers.openai.com/api/docs/guides/reasoning#phase-parameter).

81- **Compaction:** For long-running agents, use [conversation/state compaction](https://developers.openai.com/api/docs/guides/compaction) intentionally. Preserve completed actions, active assumptions, IDs, tool outcomes, unresolved blockers, and the next concrete goal.

82- **Agents SDK:** For new agentic systems, use the latest [Agents SDK](https://developers.openai.com/api/docs/guides/agents) patterns for tool orchestration, tracing, handoffs, and state management rather than rebuilding orchestration from scratch.

83- **Current date:** GPT-5.5 is aware of the current date in UTC. You don't need to add the current date to system instructions. Add explicit date or timezone context only when the application needs a business-specific timezone, policy-effective date, user-local date, or other non-UTC reference point.

84 

85## Prompting best practices

86 

87GPT-5.5 works best when prompts define the outcome and leave room for the model to choose an efficient solution path. Compared with earlier models, you can often use shorter, more outcome-oriented prompts: describe what good looks like, what constraints matter, what evidence is available, and what the final answer should contain.

88 

89Avoid carrying over every instruction from an older prompt stack. Legacy prompts often over-specify the process because earlier models needed more help staying on track. With GPT-5.5, that can add noise, narrow the model's search space, or lead to overly mechanical answers.

90 

91The patterns here are starting points. Adapt them to your product surface, tools, evals, and user experience goals.

92 

93### Personality and behavior

94 

95GPT-5.5's default style is efficient, direct, and task-oriented. This is useful for production systems: responses stay focused, behavior is easier to steer, and the model avoids unnecessary conversational padding.

96 

97For customer-facing assistants, support workflows, coaching experiences, and other conversational products, define both personality and collaboration style.

98 

99- **Personality** controls how the assistant sounds: tone, warmth, directness, formality, humor, empathy, and level of polish.

100- **Collaboration style** controls how the assistant works: when it asks questions, when it makes assumptions, how proactive it should be, how much context it gives, when it checks work, and how it handles uncertainty or risk.

101 

102Keep both short. Personality instructions should shape the user experience. Collaboration instructions should shape task behavior. Neither should replace clear goals, success criteria, tool rules, or stopping conditions.

103 

104Example personality block for a steady task-focused assistant:

105 

106```text

107# Personality

108You are a capable collaborator: approachable, steady, and direct. Assume the user is competent and acting in good faith, and respond with patience, respect, and practical helpfulness.

109 

110Prefer making progress over stopping for clarification when the request is already clear enough to attempt. Use context and reasonable assumptions to move forward. Ask for clarification only when the missing information would materially change the answer or create meaningful risk, and keep any question narrow.

111 

112Stay concise without becoming curt. Give enough context for the user to understand and trust the answer, then stop. Use examples, comparisons, or simple analogies when they make the point easier to grasp. When correcting the user or disagreeing, be candid but constructive. When an error is pointed out, acknowledge it plainly and focus on fixing it.

113 

114Match the user's tone within professional bounds. Avoid emojis and profanity by default, unless the user explicitly asks for that style or has clearly established it as appropriate for the conversation.

115```

116 

117Example personality block for an expressive collaborative assistant:

118 

119```text

120# Personality

121Adopt a vivid conversational presence: intelligent, curious, playful when appropriate, and attentive to the user's thinking. Ask good questions when the problem is blurry, then become decisive once there is enough context.

122 

123Be warm, collaborative, and polished. Conversation should feel easy and alive, but not chatty for its own sake. Offer a real point of view rather than merely mirroring the user, while staying responsive to their goals and constraints.

124 

125Be thoughtful and grounded when the task calls for synthesis or advice. State a clear recommendation when you have enough context, explain important tradeoffs, and name uncertainty without becoming evasive.

126```

127 

128For more expressive products, add warmth, curiosity, humor, or point of view explicitly, but keep the block short. Use personality to shape the experience, not to compensate for unclear goals or missing task instructions.

129 

130### Improve time to first visible token with a preamble

131 

132In streaming applications, users notice how long it takes before the first visible response appears. GPT-5.5 may spend time reasoning, planning, or preparing tool calls before emitting visible text.

133 

134For longer or tool-heavy tasks, prompt the model to start with a short preamble: a brief visible update that acknowledges the request and states the first step. This can improve perceived responsiveness without changing the underlying task.

135 

136Use this pattern when the task may take more than one step, require tool calls, or involve a long-running agent workflow.

137 

138```text

139Before any tool calls for a multi-step task, send a short user-visible update that acknowledges the request and states the first step. Keep it to one or two sentences.

140```

141 

142For coding agents that expose separate message phases, you can be more explicit:

143 

144```text

145You must always start with an intermediary update before any content in the analysis channel if the task will require calling tools. The user update should acknowledge the request and explain your first step.

146```

147 

148### Outcome-first prompts and stopping conditions

149 

150GPT-5.5 is strongest when the prompt defines the target outcome, success criteria, constraints, and available context, then lets the model choose the path.

151 

152For many tasks, describe the destination rather than every step. This gives the model room to choose the right search, tool, or reasoning strategy for the task.

153 

154Prefer this:

155 

156```text

157Resolve the customer's issue end to end.

158 

159Success means:

160- the eligibility decision is made from the available policy and account data

161- any allowed action is completed before responding

162- the final answer includes completed_actions, customer_message, and blockers

163- if evidence is missing, ask for the smallest missing field

164```

165 

166**Avoid unnecessary absolute rules.** Older prompts often use strict instructions like `ALWAYS`, `NEVER`, `must`, and `only` to control model behavior. Use those words for true invariants, such as safety rules, required output fields, or actions that should never happen. For judgment calls, such as when to search, ask for clarification, use a tool, or keep iterating, prefer decision rules instead.

167 

168Avoid this style of instruction unless every step is truly required:

169 

170```text

171First inspect A, then inspect B, then compare every field, then think through

172all possible exceptions, then decide which tool to call, then call the tool,

173then explain the entire process to the user.

174```

175 

176Add explicit stopping conditions:

177 

178```text

179Resolve the user query in the fewest useful tool loops, but do not let loop minimization outrank correctness, accessible fallback evidence, calculations, or required citation tags for factual claims.

180 

181After each result, ask: "Can I answer the user's core request now with useful evidence and citations for the factual claims?" If yes, answer.

182```

183 

184Define missing-evidence behavior:

185 

186```text

187Use the minimum evidence sufficient to answer correctly, cite it precisely, then stop.

188```

189 

190### Formatting

191 

192GPT-5.5 is highly steerable on output format and structure. Use that control when it improves comprehension or product fit.

193 

194Set `text.verbosity`, describe the expected output shape, and reserve heavier structure for cases where it improves comprehension or your product UI needs a stable artifact. The API default for `text.verbosity` is `medium`; use `low` when you prefer shorter, more concise responses.

195 

196Plain conversational formatting:

197 

198```text

199Let formatting serve comprehension. Use plain paragraphs as the default format for normal conversation, explanations, reports, documentation, and technical writeups. Keep the presentation clean and readable without making the structure feel heavier than the content.

200 

201Use headers, bold text, bullets, and numbered lists sparingly. Reach for them when the user requests them, when the answer needs clear comparison or ranking, or when the information would be harder to scan as prose. Otherwise, favor short paragraphs and natural transitions.

202 

203Respect formatting preferences from the user. If they ask for a terse answer, minimal formatting, no bullets, no headers, or a specific structure, follow that preference unless there is a strong reason not to.

204```

205 

206Add explicit audience and length guidance:

207 

208```text

209Write for a senior business audience. Keep the answer under 400 words. Use short paragraphs and only include bullets when they improve scannability. Prioritize the conclusion first, then the reasoning, then caveats.

210```

211 

212For editing, rewriting, summaries, or customer-facing messages, tell the model what to preserve before asking it to improve style. This pattern is useful when you want polish without expansion.

213 

214```text

215Preserve the requested artifact, length, structure, and genre first. Quietly improve clarity, flow, and correctness. Do not add new claims, extra sections, or a more promotional tone unless explicitly requested.

216```

217 

218### Grounding, citations, and retrieval budgets

219 

220For grounded answers, citation behavior should be part of the prompt. Define what needs support, what counts as enough evidence, and how the model should behave when evidence is missing. Absence of evidence shouldn't automatically become a factual "no." For more details and examples, see the [citation formatting guide](https://developers.openai.com/api/docs/guides/citation-formatting).

221 

222#### Add an explicit retrieval budget

223 

224Retrieval budgets are stopping rules for search. They tell the model when enough evidence is enough.

225 

226```text

227For ordinary Q&A, start with one broad search using short, discriminative keywords. If the top results contain enough citable support for the core request, answer from those results instead of searching again.

228 

229Make another retrieval call only when:

230- The top results do not answer the core question.

231- A required fact, parameter, owner, date, ID, or source is missing.

232- The user asked for exhaustive coverage, a comparison, or a comprehensive list.

233- A specific document, URL, email, meeting, record, or code artifact must be read.

234- The answer would otherwise contain an important unsupported factual claim.

235 

236Do not search again to improve phrasing, add examples, cite nonessential details, or support wording that can safely be made more generic.

237```

238 

239### Creative drafting guardrails

240 

241For drafting tasks, tell the model which claims must come from sources and which parts may be creatively written. This is especially important for slides, launch copy, customer summaries, talk tracks, leadership blurbs, and narrative framing.

242 

243```text

244For creative or generative requests such as slides, leadership blurbs, outbound copy, summaries for sharing, talk tracks, or narrative framing, distinguish source-backed facts from creative wording.

245 

246- Use retrieved or provided facts for concrete product, customer, metric, roadmap, date, capability, and competitive claims, and cite those claims.

247- Do not invent specific names, first-party data claims, metrics, roadmap status, customer outcomes, or product capabilities to make the draft sound stronger.

248- If there is little or no citable support, write a useful generic draft with placeholders or clearly labeled assumptions rather than unsupported specifics.

249```

250 

251### Frontend engineering and visual taste

252 

253For frontend work, refer to the [example instructions](https://developers.openai.com/api/docs/guides/frontend-prompt) for practical ways to steer UI quality. They cover product and user context, design-system alignment, first-screen usability, familiar controls, expected states, responsive behavior, and common generated-UI defaults to avoid, such as generic heroes, nested cards, decorative gradients, visible instructional text, and broken layouts.

254 

255### Prompt the model to check its work

256 

257Give GPT-5.5 access to tools that let it check outputs when validation is possible.

258 

259For coding agents, ask for concrete validation commands:

260 

261```text

262After making changes, run the most relevant validation available:

263- targeted unit tests for changed behavior

264- type checks or lint checks when applicable

265- build checks for affected packages

266- a minimal smoke test when full validation is too expensive

267 

268If validation cannot be run, explain why and describe the next best check.

269```

270 

271For visual artifacts, ask for inspection after rendering:

272 

273```text

274Render the artifact before finalizing. Inspect the rendered output for layout, clipping, spacing, missing content, and visual consistency. Revise until the rendered output matches the requirements.

275```

276 

277For engineering and planning tasks, make implementation plans traceable:

278 

279```text

280For implementation plans, include:

281- requirements and where each is addressed

282- named resources, files, APIs, or systems involved

283- state transitions or data flow where relevant

284- validation commands or checks

285- failure behavior

286- privacy and security considerations

287- open questions that materially affect implementation

288```

289 

290### Phase parameter

291 

292Starting with GPT-5.4, long-running or tool-heavy Responses workflows can use assistant-item `phase` values to distinguish intermediate updates from final answers. GPT-5.5 uses the same pattern.

293 

294If you use `previous_response_id`, the API preserves prior assistant state automatically. If your application manually replays assistant output items into the next request, preserve each original `phase` value and pass it back unchanged. This matters most when a response includes preambles, repeated tool calls, or a final answer after intermediate assistant updates.

295 

296```text

297If manually replaying assistant items:

298- Preserve assistant `phase` values exactly.

299- Use `phase: "commentary"` for intermediate user-visible updates.

300- Use `phase: "final_answer"` for the completed answer.

301- Do not add `phase` to user messages.

302```

303 

304### Suggested prompt structure

305 

306Use this structure as a starting point for complex prompts. Keep each section short. Add detail only where it changes behavior.

307 

308```text

309Role: [1-2 sentences defining the model's function, context, and job]

310 

311# Personality

312[tone, demeanor, and collaboration style]

313 

314# Goal

315[user-visible outcome]

316 

317# Success criteria

318[what must be true before the final answer]

319 

320# Constraints

321[policy, safety, business, evidence, and side-effect limits]

322 

323# Output

324[sections, length, and tone]

325 

326# Stop rules

327[when to retry, fallback, abstain, ask, or stop]

328```

guides/latest-model/gpt-5.6.md +0 −228 deleted

File Deleted View Diff

1latestModelInfo:

2 model: gpt-5.6-sol

3 migrationGuide: /api/docs/guides/upgrading-to-gpt-5p6-sol.md

4 promptingGuide: /api/docs/guides/prompt-guidance-gpt-5p6.md

5 

6# Using GPT-5.6

7 

8> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

9 

10## Introduction

11 

12GPT-5.6 sets a new quality and efficiency baseline for complex production workflows. GPT-5.6 is especially token-efficient and improves frontend aesthetics, including layout, visual hierarchy, and design judgment.

13 

14GPT-5.6 also introduces a new naming scheme. The `gpt-5.6` alias routes requests to `gpt-5.6-sol`, the model for flagship capability. Use `gpt-5.6-terra` for strong performance at a lower price and `gpt-5.6-luna` for efficient, high-volume workloads.

15 

16When migrating from GPT-5.5 or GPT-5.4, start with your current GPT-5.5 or GPT-5.4 reasoning setting, then test the same setting and one level lower on representative tasks. GPT-5.6 can often maintain or improve quality with fewer tokens, but the best setting depends on your workload.

17 

18<a id="what-is-new" className="scroll-mt-[110px]"></a>

19 

20## What's new

21 

22- **Programmatic Tool Calling:** GPT-5.6 can write JavaScript to call eligible tools, pass results between calls, and process intermediate outputs in a hosted runtime. Use [Programmatic Tool Calling](https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling) for bounded, tool-heavy workflows that do not require fresh model judgment between each step. Programmatic Tool Calling is ZDR-compatible with no additional container costs.

23- **Multi-agent [beta]:** [Multi-agent](https://developers.openai.com/api/docs/guides/responses-multi-agent) lets a GPT-5.6 instance coordinate multiple subagents in parallel and synthesize their results. Similar to ultra mode in Codex, this can reduce wall-clock time and improve performance for complex tasks that divide cleanly into independent workstreams. Multi-agent is available as a beta feature in the Responses API as we iterate on developer feedback.

24- **Explicit prompt caching:** GPT-5.6 lets you mark exactly which reusable prompt prefixes OpenAI caches. You can still use automatic caching in implicit mode. OpenAI bills cache writes at 1.25× the uncached input rate, while cache reads remain discounted. Learn how to [configure prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching).

25- **Persisted reasoning:** GPT-5.6 can reuse available reasoning items across turns to improve multi-turn quality and cache efficiency. Use `reasoning.context` to select the behavior. Learn how to [preserve reasoning across calls](https://developers.openai.com/api/docs/guides/reasoning#preserve-reasoning-across-calls).

26- **Max reasoning effort:** GPT-5.6 supports `max` reasoning effort for demanding tasks that need more exploration and verification. If you currently use `xhigh`, compare both settings on representative workloads.

27- **Pro mode:** GPT-5.6 can perform more model work to improve reliability on difficult tasks and return a single final answer. Enable it with `reasoning.mode: "pro"` when quality matters more than latency and token usage. Learn how to [use pro mode](https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode).

28- **Token efficiency:** GPT-5.6 reaches flagship-level performance with fewer output tokens.

29- **Frontend design:** GPT-5.6 creates more polished and usable websites and applications, with stronger layout, visual hierarchy, and design judgment.

30- **Intent understanding:** GPT-5.6 can better infer the user's underlying goal and intended level of work from context, so you often do not need to prescribe every step. Continue to provide domain context, hard constraints, approval boundaries, and success criteria. Tell the model when an important ambiguity should trigger a question.

31- **Original image detail:** GPT-5.6 preserves image dimensions with `original` or `auto` detail, except that images larger than 65,535 pixels on either side are scaled down to fit that limit. The API rejects images that still exceed the [30,000-patch limit](https://developers.openai.com/api/docs/guides/images-vision#image-input-requirements), rather than resizing them to fit it. Large images can use more input tokens and increase latency. Learn how to [choose an image detail level](https://developers.openai.com/api/docs/guides/images-vision#choose-an-image-detail-level).

32 

33## Safeguards

34 

35When using GPT-5.6 models, users may encounter safeguards that block or refuse some requests due to real-time cyber and biology misuse classifiers that are run as model outputs are generated. Other requests may take longer because generation is paused for several seconds mid-stream while these classifiers synchronously review outputs. Safeguards may occasionally intervene on legitimate work, particularly in dual-use areas where defensive and offensive activity can initially look similar.

36 

37If your application serves individual end users, send a stable, privacy-preserving `safety_identifier` with each request. See [Implement safety identifiers](https://developers.openai.com/api/docs/guides/safety-best-practices#implement-safety-identifiers) for guidance.

38 

39We are continuously evolving these safeguards so that they are robust and effective in holding up to adversarial pressure, while preserving access to legitimate work such as code review, vulnerability research, patch development, debugging, security education, and defensive testing.

40 

41 

42 

43 

44 

45## Migration quickstart

46 

47### Migrate with Codex

48 

49Codex can apply the recommended changes in this guide with the [OpenAI Docs skill](https://github.com/openai/skills/tree/main/skills/.curated/openai-docs).

50 

51```text

52$openai-docs migrate this project to the GPT-5.6 model family

53```

54 

55To use this skill in other coding agents, download it from the [OpenAI skills repository](https://github.com/openai/skills/tree/main/skills/.curated/openai-docs).

56 

57### Update API and model parameters

58 

59- Choose the target model for the workload. Use `gpt-5.6-sol` for flagship capability, `gpt-5.6-terra` for a balance of intelligence and cost, or `gpt-5.6-luna` for efficient, high-volume workloads. The `gpt-5.6` alias routes requests to `gpt-5.6-sol`.

60- Use the [Responses API](https://developers.openai.com/api/docs/guides/migrate-to-responses) for reasoning, tool-calling, and multi-turn workflows.

61- Set `reasoning.effort` intentionally. GPT-5.6 supports `none`, `low`, `medium`, `high`, `xhigh`, and `max`.

62 - If you are migrating from GPT-5.5 or GPT-5.4, preserve your current reasoning effort as the baseline, then compare one level lower.

63 - If you use `none`, keep it as your latency baseline and also test `low` when the workflow benefits from reasoning or tool use.

64 - Use `medium` as a balanced starting point and `low` for latency-sensitive workloads.

65 - Use `high` or `xhigh` when more reasoning produces a measured quality gain.

66 - Reserve `max` for the hardest quality-first workloads. Compare `max` and `xhigh` to find the best quality, latency, and cost tradeoff for your use case.

67- To use pro mode, keep your selected GPT-5.6 model and set `reasoning.mode` to `pro` in the Responses API; do not switch to a separate Pro model slug. Choose `reasoning.effort` independently. If you omit it, GPT-5.6 defaults to `medium` in both standard and pro modes. See [reasoning mode](https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode) for a request example and billing details.

68- Configure persisted reasoning based on how much prior reasoning is still relevant. GPT-5.6 models default to `all_turns`; earlier models default to `current_turn`.

69 - Omit `reasoning.context` or set it to `auto` to use `all_turns`, the GPT-5.6 default. Check the response's `reasoning.context` field to confirm the effective mode.

70 - Set `reasoning.context` to `all_turns` when the task's goals, assumptions, and priorities stay stable across turns.

71 - With `all_turns`, continue with `previous_response_id` to make reasoning from earlier responses available to the model.

72 - When managing history manually, preserve and resend previous user inputs and every response output item. For `store: false` or Zero Data Retention, replay the encrypted reasoning items that the API returns by default.

73 - Set `reasoning.context` to `current_turn` when earlier reasoning is no longer relevant.

74- Review prompt caching. You do not need to change code to keep using implicit caching. Because GPT-5.6 cache writes cost 1.25× the uncached input rate, track `cached_tokens` and `cache_write_tokens` to understand net cost. Use explicit breakpoints or `prompt_cache_options.mode: "explicit"` to avoid unnecessary writes, and replace `prompt_cache_retention` with `prompt_cache_options.ttl`.

75- To use Programmatic Tool Calling, add the `programmatic_tool_calling` tool and opt eligible tools in with `allowed_callers`. Update your application to handle `program` items, program-issued function calls, and `program_output` items while preserving each call's `call_id` and `caller` linkage. See the [Programmatic Tool Calling guide](https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling) for request and continuation examples.

76 - Benchmark the PTC-enabled workflow on representative tasks. Compare task success, final-answer completeness, required evidence, total tokens, latency, and cost. Fewer calls, turns, or intermediate outputs are improvements only when the final answer still meets the required quality bar.

77 

78## Prompting best practices

79 

80### Favor leaner prompts

81 

82Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%. Results will vary by workload, so treat these ranges as directional and validate changes on representative tasks from your own application.

83 

84To simplify prompts without losing important guidance:

85 

86- Start with a prompt and tool set that already works. Remove one group of instructions, examples, or tools at a time, then rerun the same evals.

87- State each instruction once.

88- Expose only tools relevant to the task, and keep their descriptions concise and precise.

89- Keep examples and style guidance when they encode a product requirement or correct a measured gap.

90- Track context both at the start of a run and as the conversation grows. Long sessions can amplify repeated prompt and tool content.

91 

92### Define autonomy and approval boundaries

93 

94GPT-5.6 can be proactive and persistent when carrying out multi-step tasks. Define what level of action each request authorizes so the model can continue safe, in-scope work without unnecessary pauses while stopping before external, destructive, costly, or scope-expanding actions.

95 

96A compact policy is usually sufficient:

97 

98```text

99For requests to answer, explain, review, diagnose, or plan, inspect the relevant

100materials and report the result. Do not implement changes unless the request also

101asks for them.

102 

103For requests to change, build, or fix, make the requested in-scope local changes

104and run relevant non-destructive validation without asking first.

105 

106Require confirmation for external writes, destructive actions, purchases, or a

107material expansion of scope.

108```

109 

110Name safe local actions explicitly, such as reading files, inspecting logs, editing in-scope code, and running tests. Keep the policy in one place and state each rule once. Repeating instructions such as “ask first,” “do not mutate,” or “wait for approval” can cause unnecessary approval requests for safe, expected actions.

111 

112### Set response length and style

113 

114GPT-5.6 tends to be more concise by default than GPT-5.5. When migrating, check whether broad brevity instructions such as “Be concise” or “Keep it short” are still useful. They may be unnecessary for some tasks and can sometimes make responses too brief. Keep them when they reliably produce the output your application needs.

115 

116For more consistent control across requests, use `text.verbosity` to set the default level of detail, then use the prompt for task-specific requirements.

117 

118#### Set a default with `text.verbosity`

119 

120Choose `low`, `medium`, or `high` as the default level of detail for a request. In the prompt, specify any task-specific length, structure, or required content. See [Set up `text.verbosity`](https://developers.openai.com/api/docs/guides/deployment-checklist#set-up-textverbosity) for an API example.

121 

122#### Specify what a short answer must include

123 

124When a task calls for a shorter answer, identify the information the model must preserve and the detail it can omit. For example:

125 

126```text

127Lead with the conclusion. Include the evidence needed to support it, any material

128caveat, and the next action. Omit secondary detail and repetition.

129 

130Keep all required facts, decisions, caveats, and next steps. Trim introductions,

131repetition, generic reassurance, and optional background first.

132```

133 

134This gives the model a clear priority order: preserve the content needed to complete the task, then remove lower-value detail.

135 

136#### Define the tone

137 

138Broad labels such as “friendly” or “empathetic” can be ambiguous. Describe the writing choices that define your product's tone, such as how directly to state the answer, when to acknowledge a problem, and whether reassurance or a sign-off is appropriate.

139 

140```text

141State the answer directly. If the user reports a problem, acknowledge the

142specific issue before giving the next step. Use reassurance only when it is

143relevant. Omit generic praise and unnecessary sign-offs.

144```

145 

146### Pro mode

147 

148#### Choose pro mode when quality matters most

149 

150Pro mode is a Responses API execution mode that applies more model work to a request before returning a single final answer. It can improve reliability on difficult tasks, but it increases latency and aggregates the tokens from that work in reported usage. Those tokens are billed at the selected model's standard token rates.

151 

152Use pro mode when a marginal quality improvement materially affects the outcome and the task is difficult enough to benefit, such as complex optimization, high-value coding or review, or deep analysis with clear evaluation criteria. Prefer standard mode for routine, latency-sensitive, or high-volume work, and whenever your evaluations do not show a meaningful gain from pro mode.

153 

154Reasoning mode and reasoning effort are independent. Pro mode works with any GPT-5.6 model and its supported reasoning efforts. Start with the same model and effort as your standard-mode baseline, then compare configurations on representative tasks instead of assuming that the highest effort is always the best tradeoff.

155 

156#### Configure pro mode in the API

157 

158Enable pro mode in the API request. Keep the same outcome-focused prompt you use in standard mode: state the goal, relevant context, constraints, required evidence, success criteria, and output format. You do not need to ask the model to “use pro mode,” “think harder,” or generate several candidate answers.

159 

160For example:

161 

162```text

163Review this database migration plan for failure modes that could cause data loss

164or extended downtime. For each finding, cite the relevant step, estimate impact

165and likelihood, and recommend a specific mitigation. Return the five most

166important risks in severity order.

167```

168 

169#### Compare quality and cost

170 

171Compare standard and pro modes on the same representative tasks. Measure task success, answer completeness, required evidence, total tokens, latency, and cost. Use pro mode selectively where its quality or reliability gain justifies the extra model work.

172 

173Learn more in the [reasoning mode guide](https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode).

174 

175### Programmatic Tool Calling

176 

177#### Choose Programmatic Tool Calling by task shape

178 

179Programmatic Tool Calling (PTC) works best for bounded workflows where code can process several tool results or large intermediate outputs and return a much smaller structured result. Use it for filtering, joining, ranking, deduplication, aggregation, validation, or other predictable processing.

180 

181Multiple, parallel, or dependent calls alone do not justify Programmatic Tool Calling. Prefer direct, non-PTC tool calls when:

182 

183- One call is sufficient

184- The intermediate outputs are already small

185- Each result may change the model’s next decision

186- An action requires approval

187- The final output must preserve citations or native artifacts

188 

189#### Make routing instructions task-specific

190 

191Do not rely on tool availability or generic instructions such as “use Programmatic Tool Calling efficiently” to produce the right route. When both direct and programmatic calling are available, explicitly state:

192 

193- Which bounded stage should use Programmatic Tool Calling.

194- Which tools it may call.

195- The exact output schema and required evidence.

196- Concurrency, retry, and stopping limits.

197- Which work should remain direct.

198 

199Tool descriptions should document their expected return fields, types, and error behavior. If the model cannot determine the return shape before writing the program, prefer direct tool calling so it can inspect the result before deciding how to use it.

200 

201If both routes are needed, define one clear handoff and tell the model not to switch routes or repeat completed work.

202 

203For example:

204 

205```text

206<tool_orchestration>

207Use Programmatic Tool Calling for [bounded stage] using only [eligible tools].

208Run independent calls concurrently when safe. Use only documented tool input

209and output fields.

210 

211Process and reduce the intermediate results, then emit exactly [output schema],

212including the evidence needed for the final answer.

213 

214Stop when [condition] is met. Retry transient failures at most [R] times.

215Do not repeat completed calls or perform side-effecting actions. If a required

216result is still missing, return a clear structured failure.

217 

218Use direct tool calls for [semantic judgment, approval, or final validation].

219</tool_orchestration>

220```

221 

222#### Assess the final answer

223 

224The `program_output` item and final assistant `message` are separate outputs; make sure to test both. In theory, a program can return the correct records while the message omits a required field, citation, or caveat.

225 

226Compare direct and programmatic calling on the same representative tasks. Check whether the final response is correct, complete, and includes the required evidence. Then compare total tokens, latency, cost, calls, turns, and retries. Count lower resource use as an improvement only when the response still passes your existing evals.

227 

228Learn more in the [Programmatic Tool Calling guide](https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling).

guides/prompt-guidance.md +160 −0 created

Details

1---

2latestModelInfo:

3 model: gpt-6-astra

4 migrationGuide: /api/docs/guides/latest-model/gpt-6-astra.md#migration-quickstart

5 promptingGuide: /api/docs/guides/latest-model/gpt-6-astra.md#prompting-best-practices

6---

7 

8# Using GPT-6 Astra

9 

10> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

11 

12## Introduction

13 

14GPT-6 Astra is our most intelligent model yet, with state-of-the-art performance in computer use, browsing, software engineering, science, and professional work. It excels at carrying out multistep workflows across code, browsers, and professional software. In [several evaluations](https://openai.com/index/gpt-6-astra/), Astra achieves stronger results while using substantially fewer output tokens—delivering a lower estimated API cost per task than earlier models despite its higher per-token pricing.

15 

16GPT-6 Astra is also our most aligned model yet. It excels at exercising care, respecting task boundaries, and communicating transparently. When instructions leave room for interpretation, it uses the context it has to fill in routine gaps and asks focused questions when the answer could change the outcome. It incorporates new requirements, changes course when asked, and answers side questions without losing track of the broader task.

17 

18To build with Astra, set `model` to `gpt-6-astra` in a [Responses API](https://developers.openai.com/api/docs/guides/migrate-to-responses) request.

19 

20<a id="gpt-6-astra-what-is-new" className="scroll-mt-[110px]"></a>

21 

22## What's new

23 

24- **Async tool calling:** GPT-6 Astra can continue reasoning, call other tools, or answer independent parts of a request while your application runs a tool. Set `async: true` on a function or custom tool and return its result when ready using the original `call_id`. Your application still executes the tool and manages pending work. See [Async tool calling](https://developers.openai.com/api/docs/guides/async-tool-calling) for basic usage and a developer-defined wait-tool pattern.

25- **Mid-turn steering:** Send additional user instructions while GPT-6 Astra is working, such as a correction or a change in requirements. Over a WebSocket connection, the Responses API preserves completed work and includes the update in a continuation. See [Mid-turn steering](https://developers.openai.com/api/docs/guides/steering) for the event flow and tool-result handling.

26- **Change reasoning mid-conversation while preserving cache:** Add a `configuration_update` input item to increase reasoning effort for difficult work or reduce it for routine follow-ups without rewriting the original prompt prefix. The updated reasoning effort applies until another `configuration_update` input item overrides it. See [Change reasoning mid-conversation](https://developers.openai.com/api/docs/guides/reasoning#change-reasoning-mid-conversation) for examples and compatibility.

27- **Misalignment monitoring:** As part of our [strengthened safeguards](https://openai.com/index/path-to-astra/) for GPT-6 Astra, our systems asynchronously monitor for misalignment and trigger alerts when necessary. See [Misalignment monitoring](https://developers.openai.com/api/docs/guides/safety-checks/misalignment-monitoring) for more information.

28- **Limitations:** GPT-6 Astra does not support the `none` reasoning effort. [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode) is unavailable for GPT-6 Astra with EU data residency.

29 

30GPT-6 Astra also supports the existing API capabilities available with GPT-5.6, including [computer use](https://developers.openai.com/api/docs/guides/tools-computer-use), [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs), [streaming](https://developers.openai.com/api/docs/guides/streaming-responses), [Programmatic Tool Calling](https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling), [multi-agent orchestration](https://developers.openai.com/api/docs/guides/responses-multi-agent), [prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching), [persisted reasoning](https://developers.openai.com/api/docs/guides/reasoning#preserve-reasoning-across-calls), [compaction](https://developers.openai.com/api/docs/guides/compaction), and [pro mode](https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode).

31 

32## Prompting best practices

33 

34GPT-6 Astra is more intelligent and capable than prior models like GPT-5.6 Sol, and also exhibits behavior patterns that can be optimized through prompting the model for your use case.

35 

36### GPT-6 Astra behavior

37 

38- [Initiative and follow-through](#initiative-and-follow-through) – The model is designed to be a more effective collaborator and is thus more likely to ask the user a question when additional input could materially change the result. This can cause it to stop when the user may expect it to make reasonable assumptions and persist.

39- [Instruction following](#instruction-following) – GPT-6 Astra is stronger at general instruction following than our previous models, giving you greater control over its behavior. It can be more sensitive to instructions contained in skills and other files, such as `AGENTS.md`. We **strongly recommend** auditing skills and other files accessible to your model for instructions that could influence its behavior.

40- [Personality and writing style](#personality-and-writing-style) – The model tends toward detailed, formatted responses and may use recurring phrases across sessions. Specify the writing style and structure your application needs.

41- [Subagent delegation](#subagent-delegation) – The model may delegate less often than desired for your workflow. Specify when and how much it should use subagents for parallel work.

42- [Testing and verification](#testing-and-verification) – For coding tasks, the model tends to be thorough in testing before considering a task complete. For smaller tasks, this can result in broader tests than the task requires.

43 

44### Initiative and follow-through

45 

46GPT-6 Astra is generally better than GPT-5.6 Sol and earlier models at staying coherent during long tasks. It is also more likely to ask for clarification where earlier models would make assumptions.

47 

48To encourage more autonomous work, start with this prompt:

49 

50```text

51You should infer the user's intent and task scope from the instructions and prior conversation context. Your job is to bias towards action and carry the user's intended task to completion.

52 

53When the user expresses intent to perform new work or fix an existing issue, persist until the user's intended goal is complete. Progress autonomously towards the user's goal (e.g. creating isolated worktrees / checkouts if needed, resolving merge conflicts, read-only actions, creating draft PRs etc.) unless they are clearly destructive or irreversible.

54```

55 

56When the user’s intent is unclear, the model is more likely to ask the user for clarification to proceed. Prompt the model to follow through if the user’s prompt implies authorization:

57 

58```text

59When the user's prompt indicates a request for action, such as "can you...", "I want to...", "help me..." and similar expressions, treat these as instructions to do the work and take action. Do not stop at acknowledging capability (e.g. "Yes…"), proposing a plan, or offering to continue. Do not settle for a partial or "helpful enough" solution that does not fully satisfy the user's task to save time, effort or tokens. If a task requires sustained work, complete all the necessary work until the intended outcome is fulfilled.

60```

61 

62Prompt the model to ask for approval only after preparing a concrete, reviewable result. This avoids blocking the task before the model has done the work it can, and often leads to quicker task completion.

63 

64```text

65Before asking the user clarifying questions, you should complete the work that is already authorized from context and necessary to make the proposed action concrete and reviewable. The user should be approving a concrete, reviewable result. For example, before deploying a change, writing to an external application, merging a PR or publishing a site, do all the required work first so that user approval is the final step. You don't need user permission for reversible tasks, read-only actions, reviews or fixes, or anything for which authorization is provided earlier in the session or strongly implied from the task instruction.

66 

67Do not introduce unsolicited warnings, disclaimers, approval flows, or safety/compliance checklists due to hypothetical risk.

68```

69 

70The model also likes to ask non-blocking questions as it’s working by default, so adjust these prompts to match the level of autonomy your application needs.

71 

72### Instruction following

73 

74GPT-6 Astra is better able to follow longer instructions, but can also be more sensitive to information in context. For example, unclear or conflicting guidance in a skill file may cause the model to pause and block work early. Make the priority of user instructions and skills explicit.

75 

76```text

77The user's instructions take precedence over guidelines provided in a skill. If explicit user instructions conflict with a skill's instructions, prioritize the user's instructions.

78```

79 

80Asking the model to identify the skill and instruction that caused it to pause or change direction can also be effective in providing transparency into model behavior.

81 

82```text

83If a skill causes you to ask for permission or confirmation, pause, leave requested work unfinished, or diverge from the user's intent, name and link to the exact SKILL.md file you read, quote the relevant instruction, and briefly explain how it applies. Distinguish explicit skill requirements from your interpretation of guidelines.

84```

85 

86Use this prompt to find silent and conflicting guidance when your application loads many skills and instruction files such as `AGENTS.md`.

87 

88### Personality and writing style

89 

90GPT-6 Astra tends to use lists, tables and Markdown to make responses scannable. If your application needs prose with less formatting, specify that preference.

91 

92```text

93Default to using clear, concise paragraphs, each developing one main idea. Use lists only when the information is genuinely parallel, sequential, or easier to compare, and avoid nested lists unless the hierarchy cannot be expressed clearly in prose. Use plain, simple language: familiar words, concrete examples, and precise verbs. Prefer active voice and direct statements.

94 

95Make sure to state the main point clearly and early, then develop it with the explanation and detail the reader needs. Let each sentence build on what came before. Develop the points that matter and provide enough support to be useful.

96```

97 

98For technical communication, the following prompt helps strike a balance between using clear, coherent language while remaining domain appropriate:

99 

100```text

101Use plain language over jargon, and reference technical details only to the degree that it helps illustrate an idea or your work to the user. Communicate complex concepts in a clear and cohesive manner, and calibrate your writing to the level of background knowledge assumed from the user's prompt and context.

102```

103 

104To reduce jargon and stock phrases in writing, start with this prompt:

105 

106```text

107Avoid using slop words or phrases like "Bottom Line:" in conclusions, "delve," "foster," "leverage," "it's worth noting," "importantly," "Question? Answer." or "This isn't about X. It's about Y.", "genuinely" or hyphenated compound descriptions and adjectives. Do not use concluding summary statements such as "In short:..", "The simplest mental model is:...".

108 

109State the intended action directly. Avoid adding what you won't do, what will remain unchanged, or how you'll separate or categorize results. Do not use contrastive framing such as "X, not Y" or "X—not Y" that introduces an unprompted alternative that the user didn't ask about. Avoid invented compound labels like "exact-head checks" and "editorial-row layouts", vague qualifiers, and canned transitions; use plain verbs and prepositions to state the actual relationship directly.

110```

111 

112### Subagent delegation

113 

114GPT-6 Astra is trained to be able to divide and delegate work to subagents that work in parallel. If you are implementing a multi-agent system in your harness, use the following prompt to tune how much GPT-6 Astra should delegate work:

115 

116```text

117If at any point you can parallelize work by delegating tasks to another agent (no matter if you are the root or subagent), you should do so using collaboration tools if it could save time or improve quality.

118```

119 

120Messages between agents may contain grammar or spacing errors. Use this prompt to make inter-agent messages easier to read:

121 

122```text

123Messages that you send to other agents and your final answer may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.

124```

125 

126The model tends to respond well to prompting for how and when it should delegate work to subagents, so tune this behavior to fit with your harness and multi-agent implementation.

127 

128### Testing and verification

129 

130For coding tasks, calibrate how much testing and verification a change requires. This can help avoid unnecessary tests or repeated checks for small changes.

131 

132```text

133Do not write tests for reversible, low-impact changes that mirror the implementation. If you do choose to verify your work with tests, make sure that the tests are meaningful and necessary to verify implementation.

134 

135Run tests appropriate to the change and complete required checks. Once those pass, broaden or repeat testing only when new changes, failures, or unresolved concerns justify it; otherwise, continue toward completing the task.

136```

137 

138## Migration quickstart

139 

140### Migrate with Codex

141 

142Codex can apply the recommended changes in this guide with the [OpenAI Docs skill](https://github.com/openai/codex/tree/main/codex-rs/skills/src/assets/samples/openai-docs).

143 

144```text

145$openai-docs migrate this project to GPT-6 Astra

146```

147 

148To use this skill in other coding agents, download it from the [Codex repository](https://github.com/openai/codex/tree/main/codex-rs/skills/src/assets/samples/openai-docs).

149 

150### Update API and model parameters

151 

152Set `model` to `gpt-6-astra`, then check the following:

153 

154- **Reasoning effort:** If you currently use `none` or `minimal`, start with `low` and compare results. Otherwise, preserve your current effective [reasoning effort](https://developers.openai.com/api/docs/guides/reasoning#reasoning-effort). Use `reasoning.effort` in Responses or `reasoning_effort` in Chat Completions.

155- **Tool calling:** Use the [Responses API](https://developers.openai.com/api/docs/guides/migrate-to-responses#migrating-from-chat-completions). GPT-6 Astra supports Chat Completions, but tool calling requires Responses.

156- **Unsupported parameters:** Remove `temperature`, `top_p`, and `top_logprobs`. For Chat Completions, also remove `logprobs`. For Responses, remove `message.output_text.logprobs` from `include`.

157- **Fast mode:** For EU data residency, use Standard processing. GPT-6 Astra does not support `service_tier: "fast"` or `service_tier: "priority"` with EU data residency. Fast mode for GPT-6 Astra does not include a latency SLA. See [Fast mode compatibility](https://developers.openai.com/api/docs/guides/fast-mode#is-fast-mode-compatible-with-data-residency-zero-data-retention-and-a-baa).

158- **Changing reasoning effort:** If your application changes effort between responses, use `configuration_update` items in standard, single-agent requests. Keep request-level `reasoning.effort` unchanged to preserve the prompt prefix for caching. Check the [compatibility limits](https://developers.openai.com/api/docs/guides/reasoning#change-reasoning-mid-conversation) before adopting this feature.

159- **Prompt caching:** When migrating from GPT-5.5 or earlier, replace `prompt_cache_retention` with `prompt_cache_options.ttl` set to `"30m"`. Review the [prompt caching changes](https://developers.openai.com/api/docs/guides/prompt-caching#summary-of-model-differences), including cache boundaries and cache-write billing.

160- **Unnecessary approval pauses:** If you run into issues where the model keeps asking for approval before proceeding, use the [initiative and follow-through guidance](#initiative-and-follow-through) to prompt for more autonomous execution. See the rest of [Prompting best practices](#prompting-best-practices) for guidance on instruction following, writing style, subagent delegation, and testing.

guides/prompt-guidance-gpt-5p6.md +0 −297 deleted

File Deleted View Diff

1# Prompting guidance for GPT-5.6 Sol

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5# Prompting guidance for GPT-5.6 Sol

6 

7Use this guide when adapting prompts, tool descriptions, agent instructions, or prompt stacks to GPT-5.6 Sol or the GPT-5.6 family. Pair it with the current [GPT-5.6 model guide](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6) for API details, limits, pricing, and feature availability.

8 

9GPT-5.6 works best when prompts define the outcome, important constraints, available evidence, and completion bar, then leave room for the model to choose an efficient path.

10 

11Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%. Results will vary by workload, so treat these ranges as directional and validate changes on representative tasks from your own application.

12 

13## Simplify prompts first

14 

15Start with a prompt and tool set that already works. Remove one group of instructions, examples, or tools at a time, then rerun the same evals.

16 

17Trim:

18 

19- repeated statements of the same rule;

20- repeated style or process instructions that do not change behavior;

21- examples that do not change behavior;

22- process instructions for behavior the model already performs reliably;

23- tools and tool descriptions unrelated to the task.

24 

25Keep:

26 

27- the user-visible outcome;

28- success criteria and stopping conditions;

29- safety, business, evidence, and permission constraints;

30- tool-routing rules when the route depends on context;

31- required output shape and validation requirements.

32 

33Review the remaining instructions for contradictions. GPT-5-class models follow prompt contracts closely, so conflicting rules can create more instability than missing detail.

34 

35## Outcome-first prompts and stopping conditions

36 

37Describe the destination rather than prescribing every step. GPT-5.6 can usually choose an efficient search, tool, or reasoning path when the prompt states what good looks like.

38 

39Prefer:

40 

41```text

42Resolve the customer's issue end to end.

43 

44Success means:

45- make the eligibility decision from available policy and account evidence

46- complete any allowed action before responding

47- return completed_actions, customer_message, and blockers

48- if required evidence is missing, ask for the smallest missing field

49```

50 

51Avoid unnecessary absolute rules. Use ALWAYS, NEVER, must, and only for true invariants such as safety rules, required fields, or actions that should never happen. For judgment calls, such as when to search, ask, use a tool, or keep iterating, prefer decision rules.

52 

53Preserve explicit user values. When the correct value is implicit, provide decision criteria and let the model reason from context or schema. Avoid universal defaults, keyword maps, and broad semantic shortcuts.

54 

55Add stopping conditions:

56 

57 Resolve the request in the fewest useful tool loops, but do not let loop

58 minimization outrank correctness, required evidence, calculations, or

59 required citations.

60 

61 After each result, ask whether the core request can now be answered with

62 useful evidence. If yes, answer. If required evidence is still missing,

63 name the missing fact and use the smallest useful fallback.

64 

65## Personality, collaboration, and response length

66 

67GPT-5.6 tends to be more concise by default than GPT-5.5. When migrating, check whether broad brevity instructions such as “Be concise” or “Keep it short” are still useful. They may be unnecessary for some tasks and can sometimes make responses too brief. Keep them when they reliably produce the output your application needs.

68 

69For more consistent control across requests, use `text.verbosity` to set the default level of detail, then use the prompt for task-specific requirements. Choose `low`, `medium`, or `high` as the default level of detail for a request. In the prompt, specify any task-specific length, structure, or required content. See [Set up `text.verbosity`](https://developers.openai.com/api/docs/guides/deployment-checklist#set-up-textverbosity) for an API example.

70 

71For customer-facing assistants and collaborative products, define both personality and collaboration style.

72 

73- Personality controls tone, warmth, directness, formality, humor, empathy, and polish.

74- Collaboration style controls when the model asks questions, makes assumptions, takes initiative, explains tradeoffs, checks work, and handles uncertainty.

75 

76Keep both short. Personality should shape the user experience; collaboration instructions should shape task behavior. Neither should replace clear goals, success criteria, tool rules, or stopping conditions.

77 

78When a task calls for a shorter answer, identify the information the model must preserve and the detail it can omit. For example:

79 

80 Lead with the conclusion. Include the evidence needed to support it, any material

81 caveat, and the next action. Omit secondary detail and repetition.

82 

83 Keep all required facts, decisions, caveats, and next steps. Trim introductions,

84 repetition, generic reassurance, and optional background first.

85 

86This gives the model a clear priority order: preserve the content needed to complete the task, then remove lower-value detail.

87 

88Broad labels such as “friendly” or “empathetic” can be ambiguous. Describe the writing choices that define your product's tone, such as how directly to state the answer, when to acknowledge a problem, and whether reassurance or a sign-off is appropriate.

89 

90 State the answer directly. If the user reports a problem, acknowledge the

91 specific issue before giving the next step. Use reassurance only when it is

92 relevant. Omit generic praise and unnecessary sign-offs.

93 

94Avoid blanket language rules such as “always respond in the user's language” unless that is truly the product requirement. Specify the intended output language and when it should change.

95 

96For editing, rewriting, summaries, and customer-facing drafts, tell the model what to preserve:

97 

98 Preserve the requested artifact, length, structure, genre, and factual claims

99 first. Improve clarity, flow, and correctness without adding new claims,

100 sections, or a more promotional tone unless requested.

101 

102## Define autonomy and approval boundaries

103 

104GPT-5.6 can be proactive and persistent when carrying out multi-step tasks. Define what level of action each request authorizes so the model can continue safe, in-scope work without unnecessary pauses while stopping before external, destructive, costly, or scope-expanding actions.

105 

106A compact policy is usually sufficient:

107 

108 For requests to answer, explain, review, diagnose, or plan, inspect the

109 relevant materials and report the result. Do not implement changes unless

110 the request also asks for them.

111 

112 For requests to change, build, or fix, make the requested in-scope local

113 changes and run relevant non-destructive validation without asking first.

114 

115 Require confirmation for external writes, destructive actions, purchases,

116 or a material expansion of scope.

117 

118Name safe local actions explicitly, such as reading files, inspecting logs, editing in-scope code, and running tests. Keep the policy in one place and state each rule once. Repeating instructions such as “ask first,” “do not mutate,” or “wait for approval” can cause unnecessary approval requests for safe, expected actions.

119 

120For long-running work, define the current layer of work. Distinguish research, design, implementation, review, and external coordination so the model does not silently move from one layer to another.

121 

122## Tool routing

123 

124Expose only task-relevant tools. Tool descriptions should state what the tool does, when to use it, important return fields, and error behavior.

125 

126When correctness depends on prerequisite retrieval or lookup, say so:

127 

128 Before taking an action, resolve required discovery, retrieval, and

129 validation steps. Do not skip a prerequisite because the intended final

130 state seems obvious.

131 

132When several reads are independent, parallelize them. When one result determines the next action, keep the work sequential. After parallel retrieval, synthesize before acting.

133 

134If a tool returns empty, partial, or suspiciously narrow results, try one or two meaningful fallbacks before concluding that no result exists.

135 

136## Programmatic Tool Calling

137 

138Programmatic Tool Calling (PTC) works best for bounded workflows where code can process several tool results or large intermediate outputs and return a much smaller structured result.

139 

140Multiple, parallel, or dependent calls alone do not justify Programmatic Tool Calling.

141 

142Use it for:

143 

144- filtering, joining, sorting, ranking, deduplication, and aggregation;

145- batching across many similar records;

146- repeated deterministic validation;

147- large structured results that can be reduced to a compact schema.

148 

149Prefer direct tool calls when:

150 

151- one call is sufficient;

152- intermediate outputs are already small;

153- each result may change the next decision;

154- an action requires approval;

155- the final answer must preserve citations or native artifacts;

156- the workflow requires semantic judgment between calls.

157 

158Do not rely on generic instructions such as “use Programmatic Tool Calling efficiently.” State the bounded stage, eligible tools, output schema, retry limit, stop condition, and handoff back to direct model judgment.

159 

160 Use Programmatic Tool Calling only for the bounded record-reduction stage.

161 Call only the documented read-only tools. Filter and deduplicate the

162 intermediate results, then emit exactly the required compact schema with

163 evidence fields. Retry transient failures at most twice. Use direct tool

164 calls for approval, semantic judgment, citations, and final validation.

165 

166If both routes are needed, define one clear handoff and tell the model not to switch routes or repeat completed work.

167 

168The `program_output` item and final assistant `message` are separate outputs; make sure to test both. In theory, a program can return the correct records while the message omits a required field, citation, or caveat.

169 

170Compare direct and programmatic calling on the same representative tasks. Check whether the final response is correct, complete, and includes the required evidence. Then compare total tokens, latency, cost, calls, turns, and retries. Count lower resource use as an improvement only when the response still passes your existing evals.

171 

172## Grounding, citations, and retrieval budgets

173 

174For grounded answers, citation behavior should be part of the prompt. Define what needs support, what counts as enough evidence, and how to behave when evidence is missing. Absence of evidence should not automatically become a factual “no.”

175 

176 For ordinary Q&A, start with one broad search using short, discriminative

177 keywords. If the top results contain enough support for the core request,

178 answer from those results.

179 

180 Make another retrieval call only when a required fact, owner, date, ID, or

181 source is missing; the user asked for exhaustive coverage or comparison; a

182 specific artifact must be read; or an important claim would otherwise be

183 unsupported.

184 

185 Do not search again only to improve phrasing, add examples, or support

186 nonessential detail.

187 

188For research and synthesis:

189 

190- cite only retrieved sources;

191- attach citations to the claims they support;

192- label inference separately from directly supported facts;

193- state conflicts between sources;

194- narrow the answer or report missing evidence instead of guessing.

195 

196For creative drafting, distinguish source-backed facts from creative wording. Do not invent names, metrics, dates, roadmap status, customer outcomes, or product capabilities to make a draft sound stronger.

197 

198## Long-running workflows and state

199 

200For multi-step or tool-heavy tasks, prompt for a short visible preamble before the first tool call, then sparse outcome-based updates at major phase changes. Do not ask the model to narrate routine tool calls.

201 

202 Before tool calls for a multi-step task, send a one- or two-sentence

203 user-visible update that states the first step. During the task, update only

204 when a major phase begins or a finding changes the plan. Each update should

205 state one concrete outcome and the next step.

206 

207Preserve assistant phase values when replaying history so the model can distinguish commentary from the final answer. If using previous_response_id, prior assistant state is preserved automatically. If replaying history manually, preserve each original phase value unchanged.

208 

209Compact after major milestones rather than every turn. Keep the prompt functionally consistent after compaction and treat compacted items as opaque state.

210 

211Persisted reasoning is useful when the objective, assumptions, and priorities remain stable across turns. Use current-turn behavior when earlier reasoning is no longer relevant. Do not treat persisted reasoning as an always-on optimization: stale reasoning can add tokens, increase latency, and anchor the model to an outdated approach.

212 

213Prompt caching also affects prompt construction. Keep reusable prefixes stable and avoid unnecessary churn in large system prompts. Use explicit cache breakpoints only when they improve measured cache behavior and cost for the workload.

214 

215## Reasoning effort

216 

217Establish a baseline with the current reasoning effort before changing it.

218 

219- Preserve the current GPT-5.5 or GPT-5.4 reasoning effort as the baseline.

220- Test the same setting and one level lower on representative tasks.

221- Use low for latency-sensitive work when it preserves quality.

222- Use medium as a balanced starting point.

223- Use high or xhigh only when evals show a meaningful gain.

224- Reserve max for the hardest quality-first workloads; do not recommend it globally.

225 

226Before increasing reasoning effort, check whether the prompt is missing a success criterion, dependency rule, tool-routing rule, or verification loop.

227 

228## Frontend and visual tasks

229 

230GPT-5.6 has stronger layout, visual hierarchy, and design judgment. Still provide product context, preserve the existing design system, and name the states and constraints that matter.

231 

232For incremental frontend changes:

233 

234- inspect and preserve existing design tokens, components, and patterns;

235- do not add extra features or decorative UI unless requested;

236- preserve responsive behavior and expected states;

237- render and inspect the result before finalizing.

238 

239For vision, computer use, localization, or OCR tasks where spatial precision matters, choose image detail intentionally. Use original detail for large, dense, or coordinate-sensitive images when the extra input cost and latency are justified.

240 

241## Check work before finishing

242 

243Give GPT-5.6 access to tools that can validate the output, and state what validation matters.

244 

245For coding:

246 

247```text

248After making changes, run the most relevant validation available:

249- targeted tests for changed behavior

250- type checks or lint checks when applicable

251- build checks for affected packages

252- a minimal smoke test when full validation is too expensive

253 

254If validation cannot be run, explain why and describe the next best check.

255```

256 

257For visual artifacts:

258 

259 Render the artifact before finalizing. Inspect layout, clipping, spacing,

260 missing content, and visual consistency. Revise until the rendered output

261 matches the requirements.

262 

263For implementation plans, include requirements, named resources or files, state transitions or data flow, validation checks, failure behavior, privacy or security considerations, and open questions that materially affect implementation.

264 

265## Suggested prompt structure

266 

267Use this structure as a starting point for complex prompts. Keep each section short. Add detail only where it changes behavior.

268 

269 Role: [the model's function and context]

270 

271 Personality: [tone and collaboration style]

272 

273 Goal: [user-visible outcome]

274 

275 Success criteria: [what must be true before the final answer]

276 

277 Constraints: [policy, safety, business, evidence, and side-effect limits]

278 

279 Tools: [which tools to use, when, and what not to use]

280 

281 Output: [sections, length, format, and tone]

282 

283 Stop rules: [when to retry, fallback, abstain, ask, or stop]

284 

285## Prompt migration workflow

286 

287When moving an existing application to GPT-5.6:

288 

2891. Switch the model and preserve the current reasoning effort.

2902. Run representative evals before changing the prompt.

2913. Remove obsolete scaffolding, repeated instructions, and irrelevant tools.

2924. Add only the smallest targeted instruction that fixes a measured regression.

2935. Re-run evals after each prompt or reasoning change.

294 

295Do not rewrite a working prompt stack all at once. Otherwise you cannot tell whether a behavior change came from the model, reasoning setting, prompt, tool set, or runtime.

296 

297When a prompt regresses, debug it with a small set of real traces. Identify the failure mode, find the instruction or contradiction that likely caused it, make a surgical edit, and rerun the same cases.

guides/upgrading-to-gpt-5p6-sol.md +0 −452 deleted

File Deleted View Diff

1# Upgrading to GPT-5.6 Sol

2 

3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.

4 

5# Upgrading to GPT-5.6 Sol

6 

7Use this guide when the user asks to migrate an existing OpenAI API integration, repository, prompt stack, agent, model router, or model picker to GPT-5.6 Sol or the GPT-5.6 family.

8 

9The default explicit target is `gpt-5.6-sol`. The alias `gpt-5.6` routes to Sol; use it only when the repository intentionally prefers family aliases. Do not treat every old model usage as a Sol candidate: GPT-5.6 is a family with different cost, latency, context, and quality roles.

10 

11Before changing code, use the OpenAI Docs MCP to fetch the current live GPT-5.6 model guidance:

12 

13/api/docs/guides/latest-model?model=gpt-5.6

14 

15For prompt changes, also read only the `## Prompting Best Practices` section from:

16 

17/api/docs/guides/latest-model?model=gpt-5.6#prompting-best-practices

18 

19Treat live docs as canonical for current model IDs, parameters, limits, pricing, and feature availability. This file supplies migration judgment: where to look, what can break, what to preserve, what not to adopt automatically, and how to validate the result.

20 

21## Core principle

22 

23Do not perform a blind model-string replacement.

24 

25First preserve the behavior, latency class, cost class, reasoning level, endpoint contract, tool semantics, cache behavior, and output contract of each usage site. Then make the smallest safe migration. Adopt new GPT-5.6 capabilities only when they solve a measured problem or the user explicitly asks for them.

26 

27A model upgrade alone does not authorize adding reasoning fields, changing request schemas, or rewriting tests. Only add explicit reasoning when the old effective behavior is established and omission would change behavior on GPT-5.6.

28 

29The main 5.6 migration hazards are:

30 

31- choosing Sol for workloads that were intentionally mini, nano, low-cost, or latency-sensitive;

32- inheriting 5.6's default `medium` reasoning where the old effective effort was `none`;

33- using Chat Completions with function tools without explicitly setting effective reasoning to `none`;

34- losing prompt-cache hits when a stable prefix is followed by a changing suffix;

35- increasing image or PDF input tokens because omitted or `auto` detail behaves differently;

36- applying new cache, persisted-reasoning, Pro, Programmatic Tool Calling, or multi-agent fields to routes that do not support them;

37- updating model strings but forgetting registries, allowlists, pricing metadata, capability flags, tests, and UI model pickers.

38 

39## Migration posture

40 

41Classify every usage site before editing:

42 

431. `simple Sol migration`

44 - One flagship model usage.

45 - Same endpoint and request shape can remain.

46 - Reasoning effort is explicit or its old effective value is known.

47 - No cache, vision, file, tool, or parser behavior needs implementation changes.

482. `tier-aware family migration`

49 - The repository exposes multiple model roles, model choices, fallbacks, routers, pricing data, or capability metadata.

50 - Map each role to Sol, Terra, or Luna instead of replacing everything with Sol.

513. `compatibility migration`

52 - The safe move requires parameter, endpoint, cache, state, tool-loop, or multimodal-detail changes.

53 - Make these changes only when implementation work is inside the user's requested scope. Otherwise report the exact blocker and smallest follow-up.

544. `prompt migration`

55 - The API shape can remain, but representative traces show a prompt-specific regression.

56 - Make a surgical prompt edit tied to that failure; do not rewrite a working prompt stack wholesale.

57 - When the task is to update prompting guidance, edit the directly tied prompt surface only. Do not modify runtime request code, model schemas, or tests unless the prompt change requires it.

585. `optional feature adoption`

59 - Pro mode, persisted reasoning, explicit caching, Programmatic Tool Calling, or multi-agent behavior is being added deliberately.

60 - Keep this separate from the baseline migration so its effect can be measured.

616. `leave unchanged`

62 - Historical examples, documentation about old models, snapshots, fixtures, eval baselines, comparison code, intentionally pinned fallbacks, unsupported providers, or ambiguous usages.

63 

64When intent is unclear, prefer leaving a usage unchanged and list it for confirmation over silently changing its role.

65 

66## Inventory before editing

67 

68Search for more than literal model IDs. Inventory:

69 

70- model strings, aliases, environment variables, CLI flags, config defaults, and deployment settings;

71- SDK calls to Responses, Chat Completions, Batch, or provider adapters;

72- reasoning settings, token budgets, sampling settings, and latency timeouts;

73- function tools, hosted tools, structured outputs, response parsers, and replay logic;

74- system, developer, user, and tool-description prompts tied to each usage;

75- routers, fallbacks, model allowlists, enums, regexes, validation schemas, and capability maps;

76- model picker UI, display labels, descriptions, context limits, pricing metadata, and provider catalogs;

77- prompt-cache keys, retention options, stable-prefix construction, and cache metrics;

78- image, PDF, file, OCR, and computer-use inputs;

79- tests, fixtures, snapshots, evals, analytics labels, billing tables, and docs.

80 

81When changing a default model, search every active default surface: runtime config, environment/config files, setup docs, tests, CLI defaults, and deployment examples. Update them together.

82 

83For each usage site, record:

84 

85- source model and why it appears to be used;

86- endpoint and SDK/client surface;

87- prompt surface;

88- effective reasoning effort, including defaults;

89- latency, cost, context, and quality role;

90- tools, structured outputs, caching, state replay, and multimodal inputs;

91- downstream parsers or user-visible contracts;

92- migration class and validation plan.

93 

94## Choose the target model by role

95 

96Use this as a starting map, then validate against the repository's workload:

97 

98| Existing role | Starting GPT-5.6 target | Reason |

99| ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------------- |

100| Unsuffixed GPT-5 flagship, GPT-5.5, or GPT-5.4 flagship | `gpt-5.6-sol` | Sol is the flagship-equivalent tier. |

101| Mini model, balanced lower-cost route, or medium-throughput worker | `gpt-5.6-terra` | Terra is the mini-like tier. |

102| Nano model, classification, extraction, routing, high-volume, or strict-latency route | `gpt-5.6-luna` | Luna is the nano-like tier. |

103| GPT-4.1 or GPT-4o latency-sensitive flow | Evaluate Luna and Terra first; use Sol only if quality requires it | A flagship replacement can change latency and cost materially. |

104| Reasoning-heavy or hardest quality-first flow | Start with Sol at the old effective effort | Preserve the reasoning contract before tuning. |

105| Old Pro usage | Sol plus `reasoning.mode: "pro"`, only if the user wants Pro behavior | GPT-5.6 Pro is a mode, not a separate model slug. |

106| Router, fallback, or model picker | Add the family by role | Do not collapse a multi-model design into Sol. |

107| Third-party or provider-specific model | Leave unchanged unless the user explicitly requests provider migration | Model-name similarity is not a safe mapping. |

108 

109Important limits to check in live docs:

110 

111- Sol and Terra have roughly 1.05M context and 128K maximum output.

112- Luna has a smaller 400K context and 128K maximum output.

113- Sol and Terra long-context requests above 272K input tokens can change pricing for the full request.

114 

115Do not invent prices, limits, or capability flags. Fetch them from current docs before updating a registry or UI.

116 

117For model pickers and registries, preserve existing model entries by default. Add GPT-5.6 Sol, Terra, and Luna as new options unless the user explicitly asks to replace or remove older models. Do not invent pricing, context limits, capabilities, or metadata unless confirmed from canonical docs.

118 

119If using the `gpt-5.6` alias, record the returned `response.model` during validation. Do not assume an alias and an explicit Sol slug appear identically in dashboards, rate-limit configuration, analytics, or billing metadata.

120 

121## Preserve effective reasoning before tuning

122 

123GPT-5.6 supports `none`, `low`, `medium`, `high`, `xhigh`, and `max`. If omitted, GPT-5.6 defaults to `medium`.

124 

125This is a behavioral migration hazard:

126 

127- GPT-5.5 commonly defaulted to `medium`.

128- GPT-5.4, mini, and nano usages commonly defaulted to `none`.

129- A previously omitted setting can therefore become slower, more expensive, and incompatible with Chat Completions function tools after the model swap.

130 

131For each usage:

132 

1331. If effort is explicit, preserve it for the first 5.6 run when supported.

1342. If effort is omitted and the old effective default is known, add it explicitly only when GPT-5.6's omitted default would change behavior. If both old and new omitted defaults are the same, keep it omitted.

1353. If the old effective value is unknown, do not guess. Flag it and compare the old behavior with 5.6 at the likely baseline.

1364. After the baseline passes, test the same setting and one lower on representative tasks.

1375. Use `xhigh` or `max` only for hard quality-first workloads where evals show a meaningful gain.

138 

139Do not globally recommend `max`. Before increasing effort, check whether the actual failure is a missing success criterion, dependency rule, tool-routing rule, state-replay bug, or validation loop.

140 

141Use the field shape that belongs to the endpoint.

142 

143Responses:

144 

145```json

146{

147 "model": "gpt-5.6-sol",

148 "reasoning": { "effort": "none" }

149}

150```

151 

152Chat Completions:

153 

154```json

155{

156 "model": "gpt-5.6-sol",

157 "reasoning_effort": "none"

158}

159```

160 

161## Chat Completions and function tools

162 

163This is the most important endpoint-specific check.

164 

165For GPT-5.6, function tools in Chat Completions are compatible only with effective reasoning `none`. Reasoning with tools should use the Responses API.

166 

167Because GPT-5.6 defaults to `medium`, this combination is unsafe:

168 

169```json

170{

171 "model": "gpt-5.6-luna",

172 "tools": [{ "type": "function", "function": { "...": "..." } }]

173}

174```

175 

176For a latency-sensitive Chat Completions flow that must keep function tools, explicitly preserve `none`:

177 

178```json

179{

180 "model": "gpt-5.6-luna",

181 "reasoning_effort": "none",

182 "tools": [{ "type": "function", "function": { "...": "..." } }]

183}

184```

185 

186If the application needs both reasoning and tools:

187 

188- migrate that flow to Responses when implementation changes are in scope;

189- otherwise report it as a compatibility blocker;

190- do not hide the incompatibility by removing tools, dropping required reasoning, or changing the workload's behavior without approval.

191 

192If the live API rejects the intended `none` path, treat it as a current API compatibility issue and report the exact request and error rather than inventing a workaround.

193 

194## Responses API and conversation state

195 

196Prefer Responses for reasoning, tools, multi-turn agents, and new 5.6 capabilities.

197 

198For ordinary multi-turn Responses calls, preserve the repository's existing state strategy. Do not add persisted reasoning merely because it exists.

199 

200If deliberately enabling persisted reasoning:

201 

202- use `reasoning.context: "all_turns"` only when the objective and assumptions remain stable;

203- prefer `previous_response_id` when the server can carry state;

204- when replaying manually, preserve every prior user input and every relevant output item, not only assistant text;

205- with `store: false` or ZDR, preserve and replay returned reasoning items, including `encrypted_content`;

206- use current-turn behavior when old reasoning may be stale or misleading.

207 

208For manual replay, preserve item types, IDs, call IDs, caller metadata, and assistant phase values exactly. Incomplete replay can silently reduce quality or break tool continuation.

209 

210## Prompt caching

211 

212Do not assume old cache-hit behavior survives the model swap.

213 

214GPT-5.6 implicit caching places a managed breakpoint near the latest user or tool message and no longer relies on 128-token rounding. A prompt with a large stable prefix followed by a changing suffix can therefore lose cache hits even when the stable prefix itself has not changed.

215 

216Audit:

217 

218- large reusable system/developer prompts;

219- dynamic suffixes appended to otherwise stable prompts;

220- changing timestamps, request IDs, user-specific values, or tool lists in the prefix;

221- cache keys, retention settings, and cache dashboards;

222- token accounting that assumes reads only and ignores writes.

223 

224Migration rules:

225 

226- keep reusable prefixes stable;

227- do not churn large system prompts unnecessarily;

228- compare old and new `cached_tokens`, `cache_write_tokens`, latency, and cost;

229- use explicit cache breakpoints only when a measured workload has a stable boundary that implicit caching misses;

230- do not globally convert every prompt to explicit caching;

231- do not send 5.6-only cache fields to older routes in a mixed-model system.

232 

233When old and GPT-5.6 routes share a request builder, isolate GPT-5.6-only fields instead of applying them globally.

234 

235The new top-level request shape uses `prompt_cache_options`, for example:

236 

237```json

238{

239 "prompt_cache_options": {

240 "mode": "explicit",

241 "ttl": "30m"

242 }

243}

244```

245 

246Place explicit breakpoints at the actual stable rendered boundary using `prompt_cache_breakpoint`. Preserve `prompt_cache_key` when the application already uses it. Treat the older `prompt_cache_retention` shape as deprecated and verify the live docs before rewriting it.

247 

248Cache writes cost more than ordinary uncached input, so a lower hit rate can be both slower and more expensive.

249 

250## Images, PDFs, files, and long context

251 

252GPT-5.6 can change token and latency behavior without any prompt change:

253 

254- for image inputs, omitted or `auto` image detail can preserve original dimensions;

255- for PDF/file inputs in Responses, omitted or `input_file.detail: "auto"` can use high page-image detail;

256- Chat Completions file inputs do not expose the same detail control;

257- long-context Sol and Terra requests can cross pricing thresholds;

258- Luna's smaller context can break workloads that fit in Sol or Terra.

259 

260For multimodal or long-context usages:

261 

2621. Measure input tokens and latency before and after.

2632. Make detail explicit when cost or latency matters.

2643. Resize images or use lower detail when the task does not need original spatial precision.

2654. Keep original/high detail for dense, coordinate-sensitive, OCR, localization, or visual-inspection tasks where it materially improves quality.

2665. Test worst-case context lengths, not only typical requests.

267 

268Do not claim a capability was removed based only on a missing metadata flag. Verify against current docs and a representative request.

269 

270## Structured outputs, parsers, and tool contracts

271 

272Keep output contracts explicit:

273 

274- preserve JSON schemas, required fields, enums, refusal handling, and parser expectations;

275- preserve tool names, parameter schemas, call IDs, and retry behavior;

276- keep citations, evidence fields, or native artifacts when downstream consumers require them;

277- validate that the final answer still satisfies the contract, not merely that a tool call succeeded.

278 

279Do not fix a failing migration by weakening a schema, deleting required behavior, removing routes, dropping tools, or changing business logic unless the user explicitly asked for that product change.

280 

281## Optional: Pro mode

282 

283Do not enable Pro mode during a baseline migration unless the old usage was Pro-like or the user explicitly asks for it.

284 

285GPT-5.6 Pro uses the base model with a reasoning mode:

286 

287```json

288{

289 "model": "gpt-5.6-sol",

290 "reasoning": {

291 "mode": "pro",

292 "effort": "medium"

293 }

294}

295```

296 

297Rules:

298 

299- use Responses, not Chat Completions;

300- do not search for or invent a separate `gpt-5.6-pro` slug;

301- supported Pro efforts begin at `medium`;

302- mode and effort are separate decisions;

303- compare task quality, total latency, and actual billed token usage against standard mode.

304 

305If migrating a legacy Pro slug, make the mode change explicit and evaluate it separately from ordinary Sol migration.

306 

307## Optional: Programmatic Tool Calling

308 

309Programmatic Tool Calling is not a required part of moving to GPT-5.6. Add it only when code can reduce large structured intermediate results before they return to model context.

310 

311Good candidates:

312 

313- bounded read-only filtering, joining, sorting, ranking, deduplication, and aggregation;

314- batching many similar records;

315- repeated deterministic validation;

316- map-reduce style retrieval with a compact result schema.

317 

318Poor candidates:

319 

320- one direct tool call;

321- adaptive workflows where each result changes the next decision;

322- write, approval, or side-effecting flows;

323- citation-heavy or native-artifact flows;

324- semantic judgment that should remain visible to the model.

325 

326Request-shape requirements:

327 

328```json

329{

330 "tools": [

331 { "type": "programmatic_tool_calling" },

332 {

333 "type": "function",

334 "name": "lookup_records",

335 "allowed_callers": ["programmatic"]

336 }

337 ]

338}

339```

340 

341Do not nest `programmatic_tool_calling` under another `tools` property. When enabled, the host must handle `program`, program-issued `function_call`, `function_call_output`, and `program_output` items. Preserve the original `call_id` and `caller` when returning function results.

342 

343Constrain the stage, eligible read-only tools, output schema, retry limit, and handoff back to direct judgment. Validate the final user-visible answer; a correct program result can still become an incorrect final answer.

344 

345## Optional: multi-agent beta

346 

347Do not enable multi-agent behavior during a baseline migration unless the application already has a clear parallelizable workflow and the user asks for it.

348 

349Enabling it requires:

350 

351- the `OpenAI-Beta: responses_multi_agent=v1` header;

352- `multi_agent: { "enabled": true, "max_concurrent_subagents": 3 }`;

353- handling `multi_agent_call`, `multi_agent_call_output`, and `agent_message` items;

354- executing ordinary developer-defined function calls from any agent and returning all required outputs;

355- preserving new items for replay and tracing;

356- checking incompatibilities with compaction, reasoning summaries, and tool-call limits in current docs.

357 

358Cap concurrency. Do not let a migration task create unbounded subagents, duplicate work, or finish without a final synthesis.

359 

360## Prompt migration judgment

361 

362After the model and API baseline is working, run representative traces before editing prompts. Change prompts only for measured failures.

363 

364For GPT-5.6, prefer:

365 

366- shorter, outcome-oriented prompts;

367- explicit success criteria, dependencies, stopping conditions, and completion boundaries;

368- preserved user-provided values;

369- decision criteria for implicit choices instead of universal defaults or keyword maps;

370- explicit autonomy and permission boundaries;

371- explicit tool routing, resource links, breadcrumbs, and expected tool choice;

372- staged plans, current-layer awareness, and concise handoffs for long work;

373- real validation before declaring completion.

374 

375Avoid:

376 

377- generic `be brief`, `be thorough`, or `think step by step` instructions;

378- blanket language instructions that can cause unwanted language switching;

379- repeating `ask first` until safe local work becomes blocked;

380- giant prompt rewrites that make the source of a regression impossible to identify;

381- telling the model to minimize tool loops when correctness, evidence, or required validation needs more work.

382 

383For coding or agentic migrations, add concrete preservation and verification rules:

384 

385```

386Preserve existing functionality, routes, outputs, and user-visible behavior.

387Do not delete or disable required behavior merely to make the build pass.

388Before finishing, run the relevant build, tests, type checks, render or smoke

389checks, and report the evidence.

390```

391 

392For long-running work, define the current layer: research, design, implementation, review, or external coordination. Do not let the model silently move to another layer.

393 

394## Upgrade workflow

395 

3961. Fetch current live 5.6 docs and the Prompting Best Practices section.

3972. Inventory every usage site and its adjacent prompt, config, registry, parser, and test surfaces.

3983. Classify each usage by role and migration class.

3994. Choose Sol, Terra, or Luna by the existing workload's role.

4005. Preserve the old effective reasoning effort explicitly.

4016. Run the compatibility gates:

402 - endpoint and SDK support;

403 - Chat Completions plus function tools;

404 - cache topology and cache fields;

405 - context length and long-context cost;

406 - image, PDF, and file detail;

407 - structured outputs and parsers;

408 - Responses state replay and tool continuation;

409 - mixed-model routing and unsupported new fields.

4107. Apply the smallest safe model, config, registry, and prompt changes.

4118. Do not add optional Pro, persisted reasoning, PTC, explicit caching, or multi-agent behavior unless needed and measurable.

4129. Run existing tests and representative evals.

41310. Report changed, unchanged, blocked, and confirmation-needed sites separately.

414 

415## Validation matrix

416 

417Prefer a controlled comparison:

418 

4191. old model + old prompt + old settings;

4202. GPT-5.6 target + same prompt + preserved effective reasoning;

4213. GPT-5.6 target + same prompt + one lower effort;

4224. GPT-5.6 target + the smallest prompt or API fix required by a measured failure;

4235. optional feature treatment, isolated from the baseline.

424 

425Measure what matters for the workflow:

426 

427- task success and user-visible quality;

428- structured-output validity and parser success;

429- tool choice, tool arguments, retries, loop count, and completion rate;

430- TTFT, end-to-end latency, timeout rate, and concurrency behavior;

431- input, output, reasoning, cached, and cache-write tokens;

432- cost per successful task;

433- long-context, compaction, and replay behavior;

434- image/PDF token use and visual/OCR accuracy;

435- completeness, preserved behavior, citations, and validation evidence.

436 

437For model routers and pickers, test at least one representative workload for each role. Verify that the cheapest or fastest tier is not accidentally used for quality-critical work and that Sol is not accidentally used for every workload.

438 

439## Required final report

440 

441Return:

442 

443- `Current usage inventory`: each model site, endpoint, role, prompt surface, and old effective reasoning.

444- `Target mapping`: Sol, Terra, Luna, unchanged, or confirmation-needed, with the reason.

445- `Changes made`: model strings, reasoning settings, prompts, registries, metadata, tests, and API-shape changes.

446- `Compatibility checks`: Chat Completions/tools, caching, state replay, multimodal detail, context/cost, schemas, and mixed-model routing.

447- `Prompt changes`: each surgical edit and the failure mode it addresses.

448- `Validation`: commands, evals, traces, before/after measurements, and remaining gaps.

449- `Unchanged sites`: historical, pinned, ambiguous, or intentionally role-specific usages.

450- `Blockers and open questions`: exact issue, why it is unsafe to guess, and the smallest next step.

451 

452Never say the migration is complete merely because model strings changed. It is complete only when the affected behavior and contracts have been validated or the remaining gaps are stated explicitly.