guides/upgrading-to-gpt-5p6-sol.md +0 −452 deleted
File Deleted View Diff
1# Upgrading to GPT-5.6 Sol
2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4
5# Upgrading to GPT-5.6 Sol
6
7Use this guide when the user asks to migrate an existing OpenAI API integration, repository, prompt stack, agent, model router, or model picker to GPT-5.6 Sol or the GPT-5.6 family.
8
9The default explicit target is `gpt-5.6-sol`. The alias `gpt-5.6` routes to Sol; use it only when the repository intentionally prefers family aliases. Do not treat every old model usage as a Sol candidate: GPT-5.6 is a family with different cost, latency, context, and quality roles.
10
11Before changing code, use the OpenAI Docs MCP to fetch the current live GPT-5.6 model guidance:
12
13/api/docs/guides/latest-model?model=gpt-5.6
14
15For prompt changes, also read only the `## Prompting Best Practices` section from:
16
17/api/docs/guides/latest-model?model=gpt-5.6#prompting-best-practices
18
19Treat live docs as canonical for current model IDs, parameters, limits, pricing, and feature availability. This file supplies migration judgment: where to look, what can break, what to preserve, what not to adopt automatically, and how to validate the result.
20
21## Core principle
22
23Do not perform a blind model-string replacement.
24
25First preserve the behavior, latency class, cost class, reasoning level, endpoint contract, tool semantics, cache behavior, and output contract of each usage site. Then make the smallest safe migration. Adopt new GPT-5.6 capabilities only when they solve a measured problem or the user explicitly asks for them.
26
27A model upgrade alone does not authorize adding reasoning fields, changing request schemas, or rewriting tests. Only add explicit reasoning when the old effective behavior is established and omission would change behavior on GPT-5.6.
28
29The main 5.6 migration hazards are:
30
31- choosing Sol for workloads that were intentionally mini, nano, low-cost, or latency-sensitive;
32- inheriting 5.6's default `medium` reasoning where the old effective effort was `none`;
33- using Chat Completions with function tools without explicitly setting effective reasoning to `none`;
34- losing prompt-cache hits when a stable prefix is followed by a changing suffix;
35- increasing image or PDF input tokens because omitted or `auto` detail behaves differently;
36- applying new cache, persisted-reasoning, Pro, Programmatic Tool Calling, or multi-agent fields to routes that do not support them;
37- updating model strings but forgetting registries, allowlists, pricing metadata, capability flags, tests, and UI model pickers.
38
39## Migration posture
40
41Classify every usage site before editing:
42
431. `simple Sol migration`
44 - One flagship model usage.
45 - Same endpoint and request shape can remain.
46 - Reasoning effort is explicit or its old effective value is known.
47 - No cache, vision, file, tool, or parser behavior needs implementation changes.
482. `tier-aware family migration`
49 - The repository exposes multiple model roles, model choices, fallbacks, routers, pricing data, or capability metadata.
50 - Map each role to Sol, Terra, or Luna instead of replacing everything with Sol.
513. `compatibility migration`
52 - The safe move requires parameter, endpoint, cache, state, tool-loop, or multimodal-detail changes.
53 - Make these changes only when implementation work is inside the user's requested scope. Otherwise report the exact blocker and smallest follow-up.
544. `prompt migration`
55 - The API shape can remain, but representative traces show a prompt-specific regression.
56 - Make a surgical prompt edit tied to that failure; do not rewrite a working prompt stack wholesale.
57 - When the task is to update prompting guidance, edit the directly tied prompt surface only. Do not modify runtime request code, model schemas, or tests unless the prompt change requires it.
585. `optional feature adoption`
59 - Pro mode, persisted reasoning, explicit caching, Programmatic Tool Calling, or multi-agent behavior is being added deliberately.
60 - Keep this separate from the baseline migration so its effect can be measured.
616. `leave unchanged`
62 - Historical examples, documentation about old models, snapshots, fixtures, eval baselines, comparison code, intentionally pinned fallbacks, unsupported providers, or ambiguous usages.
63
64When intent is unclear, prefer leaving a usage unchanged and list it for confirmation over silently changing its role.
65
66## Inventory before editing
67
68Search for more than literal model IDs. Inventory:
69
70- model strings, aliases, environment variables, CLI flags, config defaults, and deployment settings;
71- SDK calls to Responses, Chat Completions, Batch, or provider adapters;
72- reasoning settings, token budgets, sampling settings, and latency timeouts;
73- function tools, hosted tools, structured outputs, response parsers, and replay logic;
74- system, developer, user, and tool-description prompts tied to each usage;
75- routers, fallbacks, model allowlists, enums, regexes, validation schemas, and capability maps;
76- model picker UI, display labels, descriptions, context limits, pricing metadata, and provider catalogs;
77- prompt-cache keys, retention options, stable-prefix construction, and cache metrics;
78- image, PDF, file, OCR, and computer-use inputs;
79- tests, fixtures, snapshots, evals, analytics labels, billing tables, and docs.
80
81When changing a default model, search every active default surface: runtime config, environment/config files, setup docs, tests, CLI defaults, and deployment examples. Update them together.
82
83For each usage site, record:
84
85- source model and why it appears to be used;
86- endpoint and SDK/client surface;
87- prompt surface;
88- effective reasoning effort, including defaults;
89- latency, cost, context, and quality role;
90- tools, structured outputs, caching, state replay, and multimodal inputs;
91- downstream parsers or user-visible contracts;
92- migration class and validation plan.
93
94## Choose the target model by role
95
96Use this as a starting map, then validate against the repository's workload:
97
98| Existing role | Starting GPT-5.6 target | Reason |
99| ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------------- |
100| Unsuffixed GPT-5 flagship, GPT-5.5, or GPT-5.4 flagship | `gpt-5.6-sol` | Sol is the flagship-equivalent tier. |
101| Mini model, balanced lower-cost route, or medium-throughput worker | `gpt-5.6-terra` | Terra is the mini-like tier. |
102| Nano model, classification, extraction, routing, high-volume, or strict-latency route | `gpt-5.6-luna` | Luna is the nano-like tier. |
103| GPT-4.1 or GPT-4o latency-sensitive flow | Evaluate Luna and Terra first; use Sol only if quality requires it | A flagship replacement can change latency and cost materially. |
104| Reasoning-heavy or hardest quality-first flow | Start with Sol at the old effective effort | Preserve the reasoning contract before tuning. |
105| Old Pro usage | Sol plus `reasoning.mode: "pro"`, only if the user wants Pro behavior | GPT-5.6 Pro is a mode, not a separate model slug. |
106| Router, fallback, or model picker | Add the family by role | Do not collapse a multi-model design into Sol. |
107| Third-party or provider-specific model | Leave unchanged unless the user explicitly requests provider migration | Model-name similarity is not a safe mapping. |
108
109Important limits to check in live docs:
110
111- Sol and Terra have roughly 1.05M context and 128K maximum output.
112- Luna has a smaller 400K context and 128K maximum output.
113- Sol and Terra long-context requests above 272K input tokens can change pricing for the full request.
114
115Do not invent prices, limits, or capability flags. Fetch them from current docs before updating a registry or UI.
116
117For model pickers and registries, preserve existing model entries by default. Add GPT-5.6 Sol, Terra, and Luna as new options unless the user explicitly asks to replace or remove older models. Do not invent pricing, context limits, capabilities, or metadata unless confirmed from canonical docs.
118
119If using the `gpt-5.6` alias, record the returned `response.model` during validation. Do not assume an alias and an explicit Sol slug appear identically in dashboards, rate-limit configuration, analytics, or billing metadata.
120
121## Preserve effective reasoning before tuning
122
123GPT-5.6 supports `none`, `low`, `medium`, `high`, `xhigh`, and `max`. If omitted, GPT-5.6 defaults to `medium`.
124
125This is a behavioral migration hazard:
126
127- GPT-5.5 commonly defaulted to `medium`.
128- GPT-5.4, mini, and nano usages commonly defaulted to `none`.
129- A previously omitted setting can therefore become slower, more expensive, and incompatible with Chat Completions function tools after the model swap.
130
131For each usage:
132
1331. If effort is explicit, preserve it for the first 5.6 run when supported.
1342. If effort is omitted and the old effective default is known, add it explicitly only when GPT-5.6's omitted default would change behavior. If both old and new omitted defaults are the same, keep it omitted.
1353. If the old effective value is unknown, do not guess. Flag it and compare the old behavior with 5.6 at the likely baseline.
1364. After the baseline passes, test the same setting and one lower on representative tasks.
1375. Use `xhigh` or `max` only for hard quality-first workloads where evals show a meaningful gain.
138
139Do not globally recommend `max`. Before increasing effort, check whether the actual failure is a missing success criterion, dependency rule, tool-routing rule, state-replay bug, or validation loop.
140
141Use the field shape that belongs to the endpoint.
142
143Responses:
144
145```json
146{
147 "model": "gpt-5.6-sol",
148 "reasoning": { "effort": "none" }
149}
150```
151
152Chat Completions:
153
154```json
155{
156 "model": "gpt-5.6-sol",
157 "reasoning_effort": "none"
158}
159```
160
161## Chat Completions and function tools
162
163This is the most important endpoint-specific check.
164
165For GPT-5.6, function tools in Chat Completions are compatible only with effective reasoning `none`. Reasoning with tools should use the Responses API.
166
167Because GPT-5.6 defaults to `medium`, this combination is unsafe:
168
169```json
170{
171 "model": "gpt-5.6-luna",
172 "tools": [{ "type": "function", "function": { "...": "..." } }]
173}
174```
175
176For a latency-sensitive Chat Completions flow that must keep function tools, explicitly preserve `none`:
177
178```json
179{
180 "model": "gpt-5.6-luna",
181 "reasoning_effort": "none",
182 "tools": [{ "type": "function", "function": { "...": "..." } }]
183}
184```
185
186If the application needs both reasoning and tools:
187
188- migrate that flow to Responses when implementation changes are in scope;
189- otherwise report it as a compatibility blocker;
190- do not hide the incompatibility by removing tools, dropping required reasoning, or changing the workload's behavior without approval.
191
192If the live API rejects the intended `none` path, treat it as a current API compatibility issue and report the exact request and error rather than inventing a workaround.
193
194## Responses API and conversation state
195
196Prefer Responses for reasoning, tools, multi-turn agents, and new 5.6 capabilities.
197
198For ordinary multi-turn Responses calls, preserve the repository's existing state strategy. Do not add persisted reasoning merely because it exists.
199
200If deliberately enabling persisted reasoning:
201
202- use `reasoning.context: "all_turns"` only when the objective and assumptions remain stable;
203- prefer `previous_response_id` when the server can carry state;
204- when replaying manually, preserve every prior user input and every relevant output item, not only assistant text;
205- with `store: false` or ZDR, preserve and replay returned reasoning items, including `encrypted_content`;
206- use current-turn behavior when old reasoning may be stale or misleading.
207
208For manual replay, preserve item types, IDs, call IDs, caller metadata, and assistant phase values exactly. Incomplete replay can silently reduce quality or break tool continuation.
209
210## Prompt caching
211
212Do not assume old cache-hit behavior survives the model swap.
213
214GPT-5.6 implicit caching places a managed breakpoint near the latest user or tool message and no longer relies on 128-token rounding. A prompt with a large stable prefix followed by a changing suffix can therefore lose cache hits even when the stable prefix itself has not changed.
215
216Audit:
217
218- large reusable system/developer prompts;
219- dynamic suffixes appended to otherwise stable prompts;
220- changing timestamps, request IDs, user-specific values, or tool lists in the prefix;
221- cache keys, retention settings, and cache dashboards;
222- token accounting that assumes reads only and ignores writes.
223
224Migration rules:
225
226- keep reusable prefixes stable;
227- do not churn large system prompts unnecessarily;
228- compare old and new `cached_tokens`, `cache_write_tokens`, latency, and cost;
229- use explicit cache breakpoints only when a measured workload has a stable boundary that implicit caching misses;
230- do not globally convert every prompt to explicit caching;
231- do not send 5.6-only cache fields to older routes in a mixed-model system.
232
233When old and GPT-5.6 routes share a request builder, isolate GPT-5.6-only fields instead of applying them globally.
234
235The new top-level request shape uses `prompt_cache_options`, for example:
236
237```json
238{
239 "prompt_cache_options": {
240 "mode": "explicit",
241 "ttl": "30m"
242 }
243}
244```
245
246Place explicit breakpoints at the actual stable rendered boundary using `prompt_cache_breakpoint`. Preserve `prompt_cache_key` when the application already uses it. Treat the older `prompt_cache_retention` shape as deprecated and verify the live docs before rewriting it.
247
248Cache writes cost more than ordinary uncached input, so a lower hit rate can be both slower and more expensive.
249
250## Images, PDFs, files, and long context
251
252GPT-5.6 can change token and latency behavior without any prompt change:
253
254- for image inputs, omitted or `auto` image detail can preserve original dimensions;
255- for PDF/file inputs in Responses, omitted or `input_file.detail: "auto"` can use high page-image detail;
256- Chat Completions file inputs do not expose the same detail control;
257- long-context Sol and Terra requests can cross pricing thresholds;
258- Luna's smaller context can break workloads that fit in Sol or Terra.
259
260For multimodal or long-context usages:
261
2621. Measure input tokens and latency before and after.
2632. Make detail explicit when cost or latency matters.
2643. Resize images or use lower detail when the task does not need original spatial precision.
2654. Keep original/high detail for dense, coordinate-sensitive, OCR, localization, or visual-inspection tasks where it materially improves quality.
2665. Test worst-case context lengths, not only typical requests.
267
268Do not claim a capability was removed based only on a missing metadata flag. Verify against current docs and a representative request.
269
270## Structured outputs, parsers, and tool contracts
271
272Keep output contracts explicit:
273
274- preserve JSON schemas, required fields, enums, refusal handling, and parser expectations;
275- preserve tool names, parameter schemas, call IDs, and retry behavior;
276- keep citations, evidence fields, or native artifacts when downstream consumers require them;
277- validate that the final answer still satisfies the contract, not merely that a tool call succeeded.
278
279Do not fix a failing migration by weakening a schema, deleting required behavior, removing routes, dropping tools, or changing business logic unless the user explicitly asked for that product change.
280
281## Optional: Pro mode
282
283Do not enable Pro mode during a baseline migration unless the old usage was Pro-like or the user explicitly asks for it.
284
285GPT-5.6 Pro uses the base model with a reasoning mode:
286
287```json
288{
289 "model": "gpt-5.6-sol",
290 "reasoning": {
291 "mode": "pro",
292 "effort": "medium"
293 }
294}
295```
296
297Rules:
298
299- use Responses, not Chat Completions;
300- do not search for or invent a separate `gpt-5.6-pro` slug;
301- supported Pro efforts begin at `medium`;
302- mode and effort are separate decisions;
303- compare task quality, total latency, and actual billed token usage against standard mode.
304
305If migrating a legacy Pro slug, make the mode change explicit and evaluate it separately from ordinary Sol migration.
306
307## Optional: Programmatic Tool Calling
308
309Programmatic Tool Calling is not a required part of moving to GPT-5.6. Add it only when code can reduce large structured intermediate results before they return to model context.
310
311Good candidates:
312
313- bounded read-only filtering, joining, sorting, ranking, deduplication, and aggregation;
314- batching many similar records;
315- repeated deterministic validation;
316- map-reduce style retrieval with a compact result schema.
317
318Poor candidates:
319
320- one direct tool call;
321- adaptive workflows where each result changes the next decision;
322- write, approval, or side-effecting flows;
323- citation-heavy or native-artifact flows;
324- semantic judgment that should remain visible to the model.
325
326Request-shape requirements:
327
328```json
329{
330 "tools": [
331 { "type": "programmatic_tool_calling" },
332 {
333 "type": "function",
334 "name": "lookup_records",
335 "allowed_callers": ["programmatic"]
336 }
337 ]
338}
339```
340
341Do not nest `programmatic_tool_calling` under another `tools` property. When enabled, the host must handle `program`, program-issued `function_call`, `function_call_output`, and `program_output` items. Preserve the original `call_id` and `caller` when returning function results.
342
343Constrain the stage, eligible read-only tools, output schema, retry limit, and handoff back to direct judgment. Validate the final user-visible answer; a correct program result can still become an incorrect final answer.
344
345## Optional: multi-agent beta
346
347Do not enable multi-agent behavior during a baseline migration unless the application already has a clear parallelizable workflow and the user asks for it.
348
349Enabling it requires:
350
351- the `OpenAI-Beta: responses_multi_agent=v1` header;
352- `multi_agent: { "enabled": true, "max_concurrent_subagents": 3 }`;
353- handling `multi_agent_call`, `multi_agent_call_output`, and `agent_message` items;
354- executing ordinary developer-defined function calls from any agent and returning all required outputs;
355- preserving new items for replay and tracing;
356- checking incompatibilities with compaction, reasoning summaries, and tool-call limits in current docs.
357
358Cap concurrency. Do not let a migration task create unbounded subagents, duplicate work, or finish without a final synthesis.
359
360## Prompt migration judgment
361
362After the model and API baseline is working, run representative traces before editing prompts. Change prompts only for measured failures.
363
364For GPT-5.6, prefer:
365
366- shorter, outcome-oriented prompts;
367- explicit success criteria, dependencies, stopping conditions, and completion boundaries;
368- preserved user-provided values;
369- decision criteria for implicit choices instead of universal defaults or keyword maps;
370- explicit autonomy and permission boundaries;
371- explicit tool routing, resource links, breadcrumbs, and expected tool choice;
372- staged plans, current-layer awareness, and concise handoffs for long work;
373- real validation before declaring completion.
374
375Avoid:
376
377- generic `be brief`, `be thorough`, or `think step by step` instructions;
378- blanket language instructions that can cause unwanted language switching;
379- repeating `ask first` until safe local work becomes blocked;
380- giant prompt rewrites that make the source of a regression impossible to identify;
381- telling the model to minimize tool loops when correctness, evidence, or required validation needs more work.
382
383For coding or agentic migrations, add concrete preservation and verification rules:
384
385```
386Preserve existing functionality, routes, outputs, and user-visible behavior.
387Do not delete or disable required behavior merely to make the build pass.
388Before finishing, run the relevant build, tests, type checks, render or smoke
389checks, and report the evidence.
390```
391
392For long-running work, define the current layer: research, design, implementation, review, or external coordination. Do not let the model silently move to another layer.
393
394## Upgrade workflow
395
3961. Fetch current live 5.6 docs and the Prompting Best Practices section.
3972. Inventory every usage site and its adjacent prompt, config, registry, parser, and test surfaces.
3983. Classify each usage by role and migration class.
3994. Choose Sol, Terra, or Luna by the existing workload's role.
4005. Preserve the old effective reasoning effort explicitly.
4016. Run the compatibility gates:
402 - endpoint and SDK support;
403 - Chat Completions plus function tools;
404 - cache topology and cache fields;
405 - context length and long-context cost;
406 - image, PDF, and file detail;
407 - structured outputs and parsers;
408 - Responses state replay and tool continuation;
409 - mixed-model routing and unsupported new fields.
4107. Apply the smallest safe model, config, registry, and prompt changes.
4118. Do not add optional Pro, persisted reasoning, PTC, explicit caching, or multi-agent behavior unless needed and measurable.
4129. Run existing tests and representative evals.
41310. Report changed, unchanged, blocked, and confirmation-needed sites separately.
414
415## Validation matrix
416
417Prefer a controlled comparison:
418
4191. old model + old prompt + old settings;
4202. GPT-5.6 target + same prompt + preserved effective reasoning;
4213. GPT-5.6 target + same prompt + one lower effort;
4224. GPT-5.6 target + the smallest prompt or API fix required by a measured failure;
4235. optional feature treatment, isolated from the baseline.
424
425Measure what matters for the workflow:
426
427- task success and user-visible quality;
428- structured-output validity and parser success;
429- tool choice, tool arguments, retries, loop count, and completion rate;
430- TTFT, end-to-end latency, timeout rate, and concurrency behavior;
431- input, output, reasoning, cached, and cache-write tokens;
432- cost per successful task;
433- long-context, compaction, and replay behavior;
434- image/PDF token use and visual/OCR accuracy;
435- completeness, preserved behavior, citations, and validation evidence.
436
437For model routers and pickers, test at least one representative workload for each role. Verify that the cheapest or fastest tier is not accidentally used for quality-critical work and that Sol is not accidentally used for every workload.
438
439## Required final report
440
441Return:
442
443- `Current usage inventory`: each model site, endpoint, role, prompt surface, and old effective reasoning.
444- `Target mapping`: Sol, Terra, Luna, unchanged, or confirmation-needed, with the reason.
445- `Changes made`: model strings, reasoning settings, prompts, registries, metadata, tests, and API-shape changes.
446- `Compatibility checks`: Chat Completions/tools, caching, state replay, multimodal detail, context/cost, schemas, and mixed-model routing.
447- `Prompt changes`: each surgical edit and the failure mode it addresses.
448- `Validation`: commands, evals, traces, before/after measurements, and remaining gaps.
449- `Unchanged sites`: historical, pinned, ambiguous, or intentionally role-specific usages.
450- `Blockers and open questions`: exact issue, why it is unsafe to guess, and the smallest next step.
451
452Never say the migration is complete merely because model strings changed. It is complete only when the affected behavior and contracts have been validated or the remaining gaps are stated explicitly.