guides/prompt-guidance-gpt-5p6.md +0 −297 deleted
File Deleted View Diff
1# Prompting guidance for GPT-5.6 Sol
2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4
5# Prompting guidance for GPT-5.6 Sol
6
7Use this guide when adapting prompts, tool descriptions, agent instructions, or prompt stacks to GPT-5.6 Sol or the GPT-5.6 family. Pair it with the current [GPT-5.6 model guide](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6) for API details, limits, pricing, and feature availability.
8
9GPT-5.6 works best when prompts define the outcome, important constraints, available evidence, and completion bar, then leave room for the model to choose an efficient path.
10
11Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%. Results will vary by workload, so treat these ranges as directional and validate changes on representative tasks from your own application.
12
13## Simplify prompts first
14
15Start with a prompt and tool set that already works. Remove one group of instructions, examples, or tools at a time, then rerun the same evals.
16
17Trim:
18
19- repeated statements of the same rule;
20- repeated style or process instructions that do not change behavior;
21- examples that do not change behavior;
22- process instructions for behavior the model already performs reliably;
23- tools and tool descriptions unrelated to the task.
24
25Keep:
26
27- the user-visible outcome;
28- success criteria and stopping conditions;
29- safety, business, evidence, and permission constraints;
30- tool-routing rules when the route depends on context;
31- required output shape and validation requirements.
32
33Review the remaining instructions for contradictions. GPT-5-class models follow prompt contracts closely, so conflicting rules can create more instability than missing detail.
34
35## Outcome-first prompts and stopping conditions
36
37Describe the destination rather than prescribing every step. GPT-5.6 can usually choose an efficient search, tool, or reasoning path when the prompt states what good looks like.
38
39Prefer:
40
41```text
42Resolve the customer's issue end to end.
43
44Success means:
45- make the eligibility decision from available policy and account evidence
46- complete any allowed action before responding
47- return completed_actions, customer_message, and blockers
48- if required evidence is missing, ask for the smallest missing field
49```
50
51Avoid unnecessary absolute rules. Use ALWAYS, NEVER, must, and only for true invariants such as safety rules, required fields, or actions that should never happen. For judgment calls, such as when to search, ask, use a tool, or keep iterating, prefer decision rules.
52
53Preserve explicit user values. When the correct value is implicit, provide decision criteria and let the model reason from context or schema. Avoid universal defaults, keyword maps, and broad semantic shortcuts.
54
55Add stopping conditions:
56
57 Resolve the request in the fewest useful tool loops, but do not let loop
58 minimization outrank correctness, required evidence, calculations, or
59 required citations.
60
61 After each result, ask whether the core request can now be answered with
62 useful evidence. If yes, answer. If required evidence is still missing,
63 name the missing fact and use the smallest useful fallback.
64
65## Personality, collaboration, and response length
66
67GPT-5.6 tends to be more concise by default than GPT-5.5. When migrating, check whether broad brevity instructions such as “Be concise” or “Keep it short” are still useful. They may be unnecessary for some tasks and can sometimes make responses too brief. Keep them when they reliably produce the output your application needs.
68
69For more consistent control across requests, use `text.verbosity` to set the default level of detail, then use the prompt for task-specific requirements. Choose `low`, `medium`, or `high` as the default level of detail for a request. In the prompt, specify any task-specific length, structure, or required content. See [Set up `text.verbosity`](https://developers.openai.com/api/docs/guides/deployment-checklist#set-up-textverbosity) for an API example.
70
71For customer-facing assistants and collaborative products, define both personality and collaboration style.
72
73- Personality controls tone, warmth, directness, formality, humor, empathy, and polish.
74- Collaboration style controls when the model asks questions, makes assumptions, takes initiative, explains tradeoffs, checks work, and handles uncertainty.
75
76Keep both short. Personality should shape the user experience; collaboration instructions should shape task behavior. Neither should replace clear goals, success criteria, tool rules, or stopping conditions.
77
78When a task calls for a shorter answer, identify the information the model must preserve and the detail it can omit. For example:
79
80 Lead with the conclusion. Include the evidence needed to support it, any material
81 caveat, and the next action. Omit secondary detail and repetition.
82
83 Keep all required facts, decisions, caveats, and next steps. Trim introductions,
84 repetition, generic reassurance, and optional background first.
85
86This gives the model a clear priority order: preserve the content needed to complete the task, then remove lower-value detail.
87
88Broad labels such as “friendly” or “empathetic” can be ambiguous. Describe the writing choices that define your product's tone, such as how directly to state the answer, when to acknowledge a problem, and whether reassurance or a sign-off is appropriate.
89
90 State the answer directly. If the user reports a problem, acknowledge the
91 specific issue before giving the next step. Use reassurance only when it is
92 relevant. Omit generic praise and unnecessary sign-offs.
93
94Avoid blanket language rules such as “always respond in the user's language” unless that is truly the product requirement. Specify the intended output language and when it should change.
95
96For editing, rewriting, summaries, and customer-facing drafts, tell the model what to preserve:
97
98 Preserve the requested artifact, length, structure, genre, and factual claims
99 first. Improve clarity, flow, and correctness without adding new claims,
100 sections, or a more promotional tone unless requested.
101
102## Define autonomy and approval boundaries
103
104GPT-5.6 can be proactive and persistent when carrying out multi-step tasks. Define what level of action each request authorizes so the model can continue safe, in-scope work without unnecessary pauses while stopping before external, destructive, costly, or scope-expanding actions.
105
106A compact policy is usually sufficient:
107
108 For requests to answer, explain, review, diagnose, or plan, inspect the
109 relevant materials and report the result. Do not implement changes unless
110 the request also asks for them.
111
112 For requests to change, build, or fix, make the requested in-scope local
113 changes and run relevant non-destructive validation without asking first.
114
115 Require confirmation for external writes, destructive actions, purchases,
116 or a material expansion of scope.
117
118Name safe local actions explicitly, such as reading files, inspecting logs, editing in-scope code, and running tests. Keep the policy in one place and state each rule once. Repeating instructions such as “ask first,” “do not mutate,” or “wait for approval” can cause unnecessary approval requests for safe, expected actions.
119
120For long-running work, define the current layer of work. Distinguish research, design, implementation, review, and external coordination so the model does not silently move from one layer to another.
121
122## Tool routing
123
124Expose only task-relevant tools. Tool descriptions should state what the tool does, when to use it, important return fields, and error behavior.
125
126When correctness depends on prerequisite retrieval or lookup, say so:
127
128 Before taking an action, resolve required discovery, retrieval, and
129 validation steps. Do not skip a prerequisite because the intended final
130 state seems obvious.
131
132When several reads are independent, parallelize them. When one result determines the next action, keep the work sequential. After parallel retrieval, synthesize before acting.
133
134If a tool returns empty, partial, or suspiciously narrow results, try one or two meaningful fallbacks before concluding that no result exists.
135
136## Programmatic Tool Calling
137
138Programmatic Tool Calling (PTC) works best for bounded workflows where code can process several tool results or large intermediate outputs and return a much smaller structured result.
139
140Multiple, parallel, or dependent calls alone do not justify Programmatic Tool Calling.
141
142Use it for:
143
144- filtering, joining, sorting, ranking, deduplication, and aggregation;
145- batching across many similar records;
146- repeated deterministic validation;
147- large structured results that can be reduced to a compact schema.
148
149Prefer direct tool calls when:
150
151- one call is sufficient;
152- intermediate outputs are already small;
153- each result may change the next decision;
154- an action requires approval;
155- the final answer must preserve citations or native artifacts;
156- the workflow requires semantic judgment between calls.
157
158Do not rely on generic instructions such as “use Programmatic Tool Calling efficiently.” State the bounded stage, eligible tools, output schema, retry limit, stop condition, and handoff back to direct model judgment.
159
160 Use Programmatic Tool Calling only for the bounded record-reduction stage.
161 Call only the documented read-only tools. Filter and deduplicate the
162 intermediate results, then emit exactly the required compact schema with
163 evidence fields. Retry transient failures at most twice. Use direct tool
164 calls for approval, semantic judgment, citations, and final validation.
165
166If both routes are needed, define one clear handoff and tell the model not to switch routes or repeat completed work.
167
168The `program_output` item and final assistant `message` are separate outputs; make sure to test both. In theory, a program can return the correct records while the message omits a required field, citation, or caveat.
169
170Compare direct and programmatic calling on the same representative tasks. Check whether the final response is correct, complete, and includes the required evidence. Then compare total tokens, latency, cost, calls, turns, and retries. Count lower resource use as an improvement only when the response still passes your existing evals.
171
172## Grounding, citations, and retrieval budgets
173
174For grounded answers, citation behavior should be part of the prompt. Define what needs support, what counts as enough evidence, and how to behave when evidence is missing. Absence of evidence should not automatically become a factual “no.”
175
176 For ordinary Q&A, start with one broad search using short, discriminative
177 keywords. If the top results contain enough support for the core request,
178 answer from those results.
179
180 Make another retrieval call only when a required fact, owner, date, ID, or
181 source is missing; the user asked for exhaustive coverage or comparison; a
182 specific artifact must be read; or an important claim would otherwise be
183 unsupported.
184
185 Do not search again only to improve phrasing, add examples, or support
186 nonessential detail.
187
188For research and synthesis:
189
190- cite only retrieved sources;
191- attach citations to the claims they support;
192- label inference separately from directly supported facts;
193- state conflicts between sources;
194- narrow the answer or report missing evidence instead of guessing.
195
196For creative drafting, distinguish source-backed facts from creative wording. Do not invent names, metrics, dates, roadmap status, customer outcomes, or product capabilities to make a draft sound stronger.
197
198## Long-running workflows and state
199
200For multi-step or tool-heavy tasks, prompt for a short visible preamble before the first tool call, then sparse outcome-based updates at major phase changes. Do not ask the model to narrate routine tool calls.
201
202 Before tool calls for a multi-step task, send a one- or two-sentence
203 user-visible update that states the first step. During the task, update only
204 when a major phase begins or a finding changes the plan. Each update should
205 state one concrete outcome and the next step.
206
207Preserve assistant phase values when replaying history so the model can distinguish commentary from the final answer. If using previous_response_id, prior assistant state is preserved automatically. If replaying history manually, preserve each original phase value unchanged.
208
209Compact after major milestones rather than every turn. Keep the prompt functionally consistent after compaction and treat compacted items as opaque state.
210
211Persisted reasoning is useful when the objective, assumptions, and priorities remain stable across turns. Use current-turn behavior when earlier reasoning is no longer relevant. Do not treat persisted reasoning as an always-on optimization: stale reasoning can add tokens, increase latency, and anchor the model to an outdated approach.
212
213Prompt caching also affects prompt construction. Keep reusable prefixes stable and avoid unnecessary churn in large system prompts. Use explicit cache breakpoints only when they improve measured cache behavior and cost for the workload.
214
215## Reasoning effort
216
217Establish a baseline with the current reasoning effort before changing it.
218
219- Preserve the current GPT-5.5 or GPT-5.4 reasoning effort as the baseline.
220- Test the same setting and one level lower on representative tasks.
221- Use low for latency-sensitive work when it preserves quality.
222- Use medium as a balanced starting point.
223- Use high or xhigh only when evals show a meaningful gain.
224- Reserve max for the hardest quality-first workloads; do not recommend it globally.
225
226Before increasing reasoning effort, check whether the prompt is missing a success criterion, dependency rule, tool-routing rule, or verification loop.
227
228## Frontend and visual tasks
229
230GPT-5.6 has stronger layout, visual hierarchy, and design judgment. Still provide product context, preserve the existing design system, and name the states and constraints that matter.
231
232For incremental frontend changes:
233
234- inspect and preserve existing design tokens, components, and patterns;
235- do not add extra features or decorative UI unless requested;
236- preserve responsive behavior and expected states;
237- render and inspect the result before finalizing.
238
239For vision, computer use, localization, or OCR tasks where spatial precision matters, choose image detail intentionally. Use original detail for large, dense, or coordinate-sensitive images when the extra input cost and latency are justified.
240
241## Check work before finishing
242
243Give GPT-5.6 access to tools that can validate the output, and state what validation matters.
244
245For coding:
246
247```text
248After making changes, run the most relevant validation available:
249- targeted tests for changed behavior
250- type checks or lint checks when applicable
251- build checks for affected packages
252- a minimal smoke test when full validation is too expensive
253
254If validation cannot be run, explain why and describe the next best check.
255```
256
257For visual artifacts:
258
259 Render the artifact before finalizing. Inspect layout, clipping, spacing,
260 missing content, and visual consistency. Revise until the rendered output
261 matches the requirements.
262
263For implementation plans, include requirements, named resources or files, state transitions or data flow, validation checks, failure behavior, privacy or security considerations, and open questions that materially affect implementation.
264
265## Suggested prompt structure
266
267Use this structure as a starting point for complex prompts. Keep each section short. Add detail only where it changes behavior.
268
269 Role: [the model's function and context]
270
271 Personality: [tone and collaboration style]
272
273 Goal: [user-visible outcome]
274
275 Success criteria: [what must be true before the final answer]
276
277 Constraints: [policy, safety, business, evidence, and side-effect limits]
278
279 Tools: [which tools to use, when, and what not to use]
280
281 Output: [sections, length, format, and tone]
282
283 Stop rules: [when to retry, fallback, abstain, ask, or stop]
284
285## Prompt migration workflow
286
287When moving an existing application to GPT-5.6:
288
2891. Switch the model and preserve the current reasoning effort.
2902. Run representative evals before changing the prompt.
2913. Remove obsolete scaffolding, repeated instructions, and irrelevant tools.
2924. Add only the smallest targeted instruction that fixes a measured regression.
2935. Re-run evals after each prompt or reasoning change.
294
295Do not rewrite a working prompt stack all at once. Otherwise you cannot tell whether a behavior change came from the model, reasoning setting, prompt, tool set, or runtime.
296
297When a prompt regresses, debug it with a small set of real traces. Identify the failure mode, find the instruction or contradiction that likely caused it, make a surgical edit, and rerun the same cases.