2 2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4 4
5## Prompt caching fundamentals5## Why prompt caching matters
6 6
7Model prompts often contain repetitive content, like system prompts and common instructions. OpenAI routes API requests to servers that recently processed the same prompt, making it faster and less expensive to reuse an exact prompt prefix than to process it from scratch. Prompt Caching works automatically for eligible requests, with no code changes required. It is enabled for all recent [models](https://developers.openai.com/api/docs/models), `gpt-4o` and newer.7Prompt caching reuses work when requests share the same prompt prefix. This provides three main benefits:
8 8
9This guide describes how prompt caching works in detail, so that you can optimize your prompts for lower latency and cost.9- **Compute-efficient:** Avoid recalculating a prompt prefix that the model has already processed.
10- **Cheaper input tokens:** Pay the model's reduced cached-input rate for reused tokens, discounted up to 90%.
11- **Faster:** Reduce the time spent processing input before the response starts.
10 12
11### Caching best practices13Prompt caching is enabled by default for supported OpenAI models. Use the [Prompt Caching Dashboard](https://platform.openai.com/usage?usage_section=prompt-caching) to monitor cache read hit rates.
12 14
13Cache hits are only possible for exact prefix matches within a prompt. To realize caching benefits, place static content like instructions and examples at the beginning of your prompt, and put variable content, such as user-specific information, at the end. This also applies to images and tools, which must be identical between requests.15## What is the prompt cache?
14 16
15- Keep instructions, tools, schemas, and shared context stable. Place request-specific content after the reusable prefix.17As the model processes input tokens, it calculates intermediate key-value (KV) states. These states let the model refer back to earlier tokens while processing new input and generating a response.
16- Set [`prompt_cache_key`](https://developers.openai.com/api/reference/resources/responses/methods/create#responses-create-prompt_cache_key) on requests that share long, common prompt prefixes. Reuse the same key for those requests to help improve cache hit rates.
17- Monitor cache reads with `cached_tokens`. On GPT-5.6 and later, use `cache_write_tokens` to compare cache-write costs with later cache reads.
18 18
1919Prompt caching preserves that state for a reusable prefix. When a later request has the same prefix and finds a matching cache entry, the model can reuse the saved state instead of processing those tokens again. It still needs to process any new input to generate a new response.
20 20
21### How prompt caching works21The prompt cache stores key-value (KV) tensors, not the tokens themselves.
22 22
23By default, caching is enabled automatically for prompts that are 1,024 tokens or longer. When you make an API request, the following steps occur:
24 23
251. **Cache routing**
26 24
27 Requests are routed to a machine based on `prompt_cache_key`, with a hash of the initial prefix of the prompt as a secondary key.25Ask ChatGPT for a deeper explanation
28 26
292. **Cache lookup**
30 27
31 The system checks whether the initial portion (prefix) of your prompt exists in the cache on the selected machine.
32 28
333. **Cache hit**29OpenAI caches the model's full rendered context including OpenAI-provided instructions, [developer messages](https://developers.openai.com/api/docs/guides/prompt-engineering#message-roles-and-instruction-following), [tool definitions](https://developers.openai.com/api/docs/guides/function-calling), and [conversation history](https://developers.openai.com/api/docs/guides/conversation-state) containing [text](https://developers.openai.com/api/docs/guides/text), [images](https://developers.openai.com/api/docs/guides/images-vision), [documents](https://developers.openai.com/api/docs/guides/file-inputs), and supported [audio](https://developers.openai.com/api/docs/guides/audio).
34 30
35 If a matching prefix is found, the system uses the cached result. This decreases latency and bills those tokens at the cached-input rate.31Cache reuse requires the entire rendered prefix to match. If content or a relevant setting changes before a breakpoint, the prefix after that change cannot match the existing cache entry.
36 32
374. **Cache miss**33### Which settings affect the cached prefix?
38 34
39 If no matching prefix is found, the system processes your full prompt. When automatic caching is enabled, it may write an eligible prefix to cache on that machine for future requests.
40 35
41For GPT-5.6 and later, 1,024 tokens is a strict minimum. For earlier models,
42 the minimum varies by model from 1,024 to 2,048 tokens, so prompts just above
43 1,024 tokens may not cache consistently.
44 36
45### How caching differs by model37Changing a request does not necessarily discard an existing cache entry. What matters is whether a subsequent request has the same prefix and can find an eligible matching breakpoint. The main settings to check are:
46 38
47| Behavior | GPT-5.6 and later | Earlier models |39| Setting | Impact |
48| -------------------------- | ------------------------------------------------------- | ------------------------------------------------------------------- |40| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
49| Cache matching | Exact matching at eligible cache breakpoints | Automatic best-effort reuse of matching prefixes |41| [`model`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20model%20%3E%20%28schema%29) | A different model can use different weights and caching behavior. |
50| Explicit cache breakpoints | Supported. Implicit caching is also available. | Not supported. Caching is automatic. |42| [`tools`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20tools%20%3E%20%28schema%29) | Changes tool names, descriptions, schemas, ordering, or tool-specific instructions. |
51| Minimum cacheable prefix | 1,024 tokens | 1,024 to 2,048 tokens, depending on the model |43| [`parallel_tool_calls`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20parallel_tool_calls%20%3E%20%28schema%29) | Can change instructions about calling multiple tools in one turn. |
52| Cache write charges | 1.25× the uncached input token rate | No additional cache-write fee |44| [`text.format`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20text%20%3E%20%28schema%29) ([Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs)) | Adds output-format instructions and the requested schema. |
53| Cache lifetime | 30-minute exact TTL set with `prompt_cache_options.ttl` | Model-dependent maximum retention set with `prompt_cache_retention` |45| [`reasoning.effort`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20reasoning%20%3E%20%28schema%29) | Can change model-side reasoning instructions. |
46| [`text.verbosity`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20text%20%3E%20%28schema%29) | Can change instructions about response detail. |
47| [`context_management`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20context_management%20%3E%20%28schema%29) ([Compaction](https://developers.openai.com/api/docs/guides/compaction)) | Replaces earlier conversation content with a compacted context that can prevent reuse from the first changed token onward. |
54 48
55For GPT-5.6 and later models, see [Prompt caching for GPT-5.6 and later models](#prompt-caching-for-gpt-56-and-later-models). For earlier models, see [Prompt caching for earlier models](#prompt-caching-for-earlier-models).
56 49
57## Prompt caching for GPT-5.6 and later models
58 50
59GPT-5.6 and later model families cache exact prompt prefixes at cache breakpoints. By default, the service places an implicit breakpoint at the latest user or tool message. Unlike earlier models, it does not automatically fall back to the longest matching unmarked prefix before that breakpoint.
60 51
61To improve cache reuse, identify the prompt content that stays the same across requests. Then choose a breakpoint that ends after that content and use a consistent `prompt_cache_key`.
62 52
63### How cache breakpoints work53## How caching works
64 54
65A cache breakpoint marks the end of a reusable prompt prefix. The prefix includes the marked content block and all prompt content rendered before it. Content after the breakpoint can change without invalidating that prefix.55A **cache breakpoint** marks the end of a prompt prefix that OpenAI can save to the cache and reuse in later requests. The first request writes an eligible prefix to the cache and a later request looks for the longest matching cached prefix available, working backward through eligible breakpoints until it finds a match.
66 56
67For a prefix to be eligible for caching, it must contain at least 1,024 tokens through the breakpoint. The minimum applies to the complete rendered prefix, not just the marked content block.57A prompt prefix must meet the model's **minimum cacheable token length** before it can be cached. Tokens in the OpenAI-provided hidden system content do not count toward this minimum. The minimum cacheable prompt length is 1,024 tokens for GPT-5.6 and later and 2,048 tokens for models older than GPT-5.6. You may occasionally get cache hits below 2,048 tokens for some earlier models. See the [model comparison](#summary-of-model-differences) for other differences.
68 58
69**Cache writes and cache reads**59After the minimum cacheable token length, you can choose where to place cache breakpoints explicitly, or let OpenAI choose their locations implicitly. The available options depend on the model.
70 60
71A cache write creates an entry for an eligible prompt prefix. A cache read reuses an entry that an earlier request wrote.
72 61
731. The first request writes an eligible prefix at a cache breakpoint.
742. A later request can read that prefix when the content through an eligible breakpoint matches the earlier cache entry and the two requests share the same `prompt_cache_key`.
753. A change before the breakpoint changes the prefix and will prevent a cache hit.
764. A change after the breakpoint does not invalidate the earlier cached prefix.
77 62
78Repeated prompt content alone does not guarantee a cache hit. If no matching entry was written at an eligible breakpoint, the system cannot read that prefix from the cache.63### GPT-5.6 and later
79 64
80**When the default breakpoint works**
81 65
82Implicit caching works well when a conversation grows by appending new messages and the earlier conversation history stays the same.
83 66
84```text67For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate. It is worth incurring this charge when a prefix will be reused, because subsequent reads cost only 0.1× that rate. Writing a prefix once and fully reusing it once costs 1.35× its ordinary input cost, compared with 2× for processing it twice without caching. The savings grow with each additional cache read: across ten requests, one write and nine full reads cost 2.15×, compared with 10× without caching.
85Request 1: Instructions → User message 1 [implicit breakpoint]
86Request 2: Instructions → User message 1 → Assistant message 1 → User message 2 [implicit breakpoint]
87```
88 68
89The first request can write the prefix through user message 1. On the next request, that earlier breakpoint can provide a cache read. The newly appended content can then be written at the latest implicit breakpoint.69Both implicit and explicit caching are supported, where explicit caching gives you more control over which context is written to cache.
90 70
91**When changing content prevents reuse**71**Explicit mode:** You choose where to place cache breakpoints based on your context management.
92 72
93Some applications send separate requests that share the same instructions but have different timestamps and user messages. Unlike successive turns in a conversation, these requests do not share conversation history.73- Set `prompt_cache_options.mode` to `explicit` to use only developer-selected breakpoints and mark each desired breakpoint by adding `prompt_cache_breakpoint: { "mode": "explicit" }` to a supported content block inside an input message.
74- When no explicit breakpoints are placed, the request does not use prompt caching or create cache writes.
75- Explicit-only mode lets you choose where cache writes end. Content after the last selected breakpoint is processed at the uncached input-token rate without a cache-write charge, so you can avoid writing changing content that is unlikely to be reused.
76- Multiple explicit breakpoints can preserve prefixes that change at different rates. Each request can create up to four cache writes.
77- For cache reads, OpenAI considers up to the latest 50 breakpoints in the conversation and reuses the longest matching cached prefix.
94 78
95```text79Top-level `instructions` cannot contain an explicit breakpoint. To mark reusable developer instructions, place them in an `input_text` block inside a developer message.
96Request 1: Stable instructions → Timestamp 1 → User message 1 [implicit breakpoint]
97Request 2: Stable instructions → Timestamp 2 → User message 2 [implicit breakpoint]
98```
99 80
100The first request writes a prefix that includes timestamp 1 and user message 1. On the second request, timestamp 2 and user message 2 change the prefix at the breakpoint. If no earlier matching entry exists, `cached_tokens` can be `0` and the service can write the changing prefix again.81**Implicit mode:** OpenAI chooses breakpoint locations out of the box that work well for most use cases.
101 82
102Add an explicit breakpoint at the end of the stable content to make that content reusable:83- When `prompt_cache_options.mode` is `implicit`, OpenAI places a breakpoint at the end of the latest eligible message.
84- You can add explicit breakpoints without turning off the implicit breakpoint; an implicit breakpoint uses one of the four cache write slots to leave three usable explicit cache write slots.
85- The implicit breakpoint creates a cache write through the latest eligible message.
103 86
104```text
105Stable instructions [explicit breakpoint] → Timestamp → User message
106```
107 87
108The first request writes the stable prefix. Later requests with the same prefix and `prompt_cache_key` can read that entry, even when the timestamp and user message change.
109 88
110### Choose a caching mode
111 89
112Use `prompt_cache_options.mode` to set the request-wide caching policy.
113 90
114**Implicit caching**
115 91
116- `implicit` is the default. OpenAI places a cache breakpoint on the latest user or tool message and also uses any explicit breakpoints you provide.
117- Use implicit caching when the prompt grows by appending reusable content. Earlier eligible breakpoints can provide cache reads, while the latest message creates a new checkpoint for future requests.
118 92
119**Explicit breakpoints with implicit caching**93### Earlier models
94
95
96
97Only implicit caching is supported. OpenAI places implicit breakpoints at [model-dependent intervals](#summary-of-model-differences), counted from the beginning of the hidden OpenAI system message. Only breakpoints at or beyond the minimum cacheable length (counted from the end of the hidden context) are eligible.
98
99Reported `cached_tokens` is calculated by subtracting the hidden system tokens from the last matched breakpoint, then rounding down to the nearest multiple of 128.
100
101
102
103
104
105## Cache lifetime
106
107Cache entries are not stored indefinitely. A later request can reuse a cached prefix only while its entry remains available, and reusing the prefix refreshes its lifetime without another cache-write charge. The lifetime and retention settings [depend on the model](#summary-of-model-differences).
108
109<a id="prompt-cache-retention"></a>
110
111
112
113### GPT-5.6 and later
114
115
116
117Use `prompt_cache_options.ttl` to control the minimum cache lifetime. The only supported value, `30m`, is also the default. A cached prefix remains eligible for reuse for 30 minutes after its most recent write or reuse, though OpenAI may retain it longer.
118
119
120
121
122
123<a id="extended-prompt-cache-retention"></a>
124
125
126
127### Earlier models
128
129
130
131Use `prompt_cache_retention`, with supported values that depend on the model:
132
133- `in_memory`: Entries typically remain active for around 5 to 10 minutes of inactivity, up to one hour.
134- `24h`: Extended retention typically keeps entries available for around 30 minutes and can retain them for up to 24 hours.
135
136**Retention defaults and Zero Data Retention**
137
138Prompt caching may store encrypted key/value tensors in GPU-local storage as application state. For models that support both `in_memory` and `24h`, the default depends on your organization's data retention policy:
139
140- Organizations _without_ Zero Data Retention enabled default to `24h`.
141- Organizations _with_ Zero Data Retention enabled default to `in_memory`.
142
143Verify the available retention policies for your model and organization before selecting a value.
144
145
146
147
148
149<a id="where-caching-happens-and-how-long-it-lasts"></a>
150
151<a id="cache-location-and-duration"></a>
152
153<a id="cache-location-and-lifetime"></a>
154
155## Cache location
156
157Cached states live on individual machines, where traffic above 15 requests per minute can lead to overflow routing. A request can reuse a cached prefix only if it reaches a machine holding a matching entry that has not expired. Routing requests to the right machine is therefore important for cache reuse.
158
159Caches are not shared across organizations and cannot be reused across [regional processing boundaries](https://developers.openai.com/api/docs/guides/your-data#data-residency-controls).
160
161OpenAI handles routing automatically. Within an organization and processing region, routing for a given model depends on:
162
163- Current machine load and available capacity.
164- A hash of the initial tokens after the hidden OpenAI content, including tool definitions when present. The number of tokens hashed varies by model.
165- The optionally supplied [`prompt_cache_key`](#prompt-cache-keys) that controls grouping and distribution during higher-volume traffic, to mitigate request overflow to other machines and, therefore, cache misses.
166
167<a id="prompt-cache-keys"></a>
168
169
170
171### Prompt cache keys
172
173
174
175When traffic exceeds a machine's available capacity, requests may overflow to another machine. If that machine does not have a matching cache entry, the initial overflow request incurs a cache miss.
176
177Set [`prompt_cache_key`](https://developers.openai.com/api/reference/resources/responses/methods/create#%28resource%29%20responses%20%3E%20%28method%29%20create%20%3E%20%28params%29%200.non_streaming%20%3E%20%28param%29%20prompt_cache_key%20%3E%20%28schema%29) to help requests with the same prefix reach the same cache. Keys influence routing; they do not pin requests to a machine or guarantee a cache read hit. See [how to tune prompt cache keys](#prompt-cache-key-best-practices).
178
179
180
181
182
183<a id="model-differences-at-a-glance"></a>
120 184
121- You can add an explicit breakpoint without changing the default caching mode. This lets requests read a stable prefix while the implicit breakpoint continues to cache the latest eligible message.185## Summary of model differences
122- This approach is useful when both the shared prefix and the growing conversation history are likely to be reused. However, the latest implicit breakpoint can still write a changing suffix to the cache.
123 186
124**Explicit-only caching**187| Behavior | GPT-5.6 and later | GPT-5.5 and GPT-5.5 Pro | Other earlier models |
188| -------------------------- | ------------------------------------------------------- | ------------------------------------------------------------------ | ------------------------------------------------------------------------------- |
189| Implicit breakpoints | At the end of the latest eligible user or tool message. | Spaced at regular 2,048-token intervals. | Spaced at regular, model-dependent intervals. |
190| Explicit breakpoints | Supported | Not supported | Not supported |
191| Minimum cacheable prefix | 1,024 visible input tokens | 2,048 visible input tokens; some models may cache shorter prefixes | 2,048 visible input tokens; some models may cache shorter prefixes |
192| Cached-token reporting | Exact eligible boundary, excluding hidden tokens | Excludes hidden tokens and rounds down to a multiple of 128 | Excludes hidden tokens and rounds down to a multiple of 128 |
193| Cache read charge | 0.1× the uncached input-token rate | Model-dependent cached-input rate | Model-dependent cached-input rate |
194| Cache write charge | 1.25× the uncached input-token rate | No additional cache-write charge | No additional cache-write charge |
195| Cache lifetime control | `prompt_cache_options.ttl` | `prompt_cache_retention` | `prompt_cache_retention` |
196| Supported retention values | `"30m"` | `"24h"` only | `"in_memory"` or `"24h"`<sup>[\*](#extended-retention-models)</sup> |
197| Cache lifetime | At least 30 minutes after the latest write or reuse | Typically around 30 minutes, up to 24 hours | Typically 5 to 10 minutes inactive for `in_memory`, or up to 24 hours for `24h` |
125 198
126- Set `prompt_cache_options.mode` to `explicit` to disable the implicit breakpoint. Only explicit breakpoints are used for cache reads and writes.199<a id="extended-retention-models"></a>
127- Use explicit-only mode when the prompt has a stable prefix followed by request-specific content that is unlikely to be reused. This caches the reusable prefix without creating a new cache write for the changing suffix.
128- Adding an explicit breakpoint does not automatically switch a request to explicit-only mode. If you set `mode` to `explicit` but provide no explicit breakpoints, the request does not use prompt caching or incur cache-write charges.
129 200
130### Add explicit cache breakpoints
131 201
132Add `prompt_cache_breakpoint: { "mode": "explicit" }` to the last supported content block in the reusable prefix. The breakpoint includes that block and all prompt content rendered before it.
133 202
134The following examples are abbreviated to show the request shape. In a real request, the rendered prefix through the marked breakpoint must contain at least 1,024 tokens.
135 203
204\* Extended retention is supported by `gpt-5.5`, `gpt-5.5-pro`, `gpt-5.4`, `gpt-5.2`, `gpt-5.1-codex-max`, `gpt-5.1`, `gpt-5.1-codex`, `gpt-5.1-codex-mini`, `gpt-5.1-chat-latest`, `gpt-5`, `gpt-5-codex`, and `gpt-4.1`.
136 205
137 206
138Responses API
139 207
140 208
141 This request places an explicit breakpoint after stable developer instructions. Explicit-only mode prevents the changing user message from creating an additional implicit cache write.209<a id="best-practices"></a>
210
211## How to optimize prompt caching
212
213Focus on [preserving conversation history](#preserve-conversation-history), [keeping tool definitions stable](#tools), and understanding the three main cache controls. Use [`prompt_cache_options.mode` and `prompt_cache_breakpoint`](#choose-a-caching-mode) to choose where caching occurs, and [`prompt_cache_key`](#prompt-cache-key-best-practices) to help related requests reach the same cache.
214
215
216
217Ask ChatGPT to optimize my prompt caching
218
219
220
221<a id="preserve-conversation-history"></a>
222
223
224
225### Preserve conversation history
226
227
228
229In multi-turn applications, reusing the growing conversation history can save more input tokens than caching only the initial instructions. Preserve earlier messages and tool results so later turns can reuse the full shared prefix.
230
231- **Keep the prefix stable.** Put stable developer instructions and shared reference material first. If developer instructions or shared material contain timestamps, user-specific content, or other dynamic content, place those at the end rather than the beginning, or move them into later conversation messages.
232- **Preserve conversation history.** Append new messages rather than rewriting earlier turns. Summarization, compaction, or context truncation can change the prefix and reset cache reuse.
233
234Keep changing content after the breakpoint
142 235
143```json236```json
144{237{
145 "model": "gpt-5.6",238 "model": "gpt-5.6",
146 "prompt_cache_key": "support:knowledge-base-v1",239 "reasoning": { "effort": "low", "context": "all_turns" },
147 "prompt_cache_options": {240 "text": { "verbosity": "medium" },
148 "mode": "explicit"241 "prompt_cache_options": { "mode": "explicit" },
149 },
150 "input": [242 "input": [
151 {243 {
152 "type": "message",
153 "role": "developer",244 "role": "developer",
154 "content": [245 "content": [
155 {246 {
156 "type": "input_text",247 "type": "input_text",
157 "text": "Follow the shared support policies and reference material...",248 "text": "Stable instructions and shared reference material...",
158 "prompt_cache_breakpoint": {249 "prompt_cache_breakpoint": { "mode": "explicit" }
159 "mode": "explicit"
160 }
161 }250 }
162 ]251 ]
163 },252 },
164 {253 {
165 "type": "message",254 "role": "developer",
166 "role": "user",255 "content": "Dynamic developer instructions, such as user-specific content and timestamps..."
167 "content": [256 },
168 {257 {
169 "type": "input_text",258 "role": "user",
170 "text": "Where is order 1234?"259 "content": "The user's current question..."
171 }
172 ]
173 }260 }
174 ]261 ]
175}262}
176```263```
177 264
178 265
179 Top-level `instructions` cannot contain a `prompt_cache_breakpoint`. To mark reusable developer instructions, place them in an `input_text` block inside a developer message, as shown above.
180 266
181 267
182 268
183 269
270<a id="tools"></a>
184 271
185 272
186Chat Completions API
187 273
274### Manage tools with append-only updates
188 275
189 This request marks the system-message prefix. Explicit-only mode limits cache reads and writes to the marked stable content.
190 276
191```json
192{
193 "model": "gpt-5.6",
194 "prompt_cache_key": "support:knowledge-base-v1",
195 "prompt_cache_options": {
196 "mode": "explicit"
197 },
198 "messages": [
199 {
200 "role": "system",
201 "content": [
202 {
203 "type": "text",
204 "text": "You are a support assistant. Follow the shared policies...",
205 "prompt_cache_breakpoint": {
206 "mode": "explicit"
207 }
208 }
209 ]
210 },
211 {
212 "role": "user",
213 "content": "What should I do next?"
214 }
215 ]
216}
217```
218 277
278When the tools your application needs vary between requests, change which tools are callable while keeping their definitions stable to preserve reusable prefixes.
219 279
280- **Keep tools consistent.** Preserve tool definitions, ordering, and schemas.
281- **Disable tool use for a request.** Set [`tool_choice`](https://developers.openai.com/api/docs/guides/function-calling#tool-choice) to `"none"` instead of removing the tool definitions.
282- **Enable only selected tools.** Use [`allowed_tools`](https://developers.openai.com/api/docs/guides/function-calling#tool-choice) to restrict which tools are callable while keeping the supplied `tools` list stable.
283- **Load tools when needed.** Use [tool search](https://developers.openai.com/api/docs/guides/tools-tool-search) with `defer_loading: true` to reduce input tokens spent on tool definitions in early requests of multi-turn threads. Discovered tools are appended at the end of context, preserving earlier reusable content.
284- **Preserve tool-loading history.** Use a developer-role [`additional_tools` input item](https://developers.openai.com/api/docs/guides/tools-tool-search#add-tools-at-a-specific-point-in-the-input) to add tools during a thread according to your application's logic.
220 285
221To combine an explicit breakpoint with the default implicit breakpoint, omit `prompt_cache_options.mode` or set it to `implicit`.
222 286
223**Supported content blocks**
224 287
225The Responses API supports breakpoints on `input_text`, `input_image`, and `input_file` blocks. The Chat Completions API supports them on `text`, `image_url`, `input_audio`, `file`, and `refusal` blocks.
226 288
227Only `explicit` is valid for `prompt_cache_breakpoint.mode`. A marker on an unsupported or non-cacheable block returns a `400 invalid_request_error`.
228 289
229Tool definitions, structured output schemas, messages, images, and files can contribute to the rendered prefix. Keep the content, order, and relevant settings identical across requests that should share a cache.290<a id="choose-a-caching-mode"></a>
230 291
231### Use multiple cache breakpoints
232 292
233Use multiple explicit breakpoints when parts of a prompt change at different rates. For example, shared instructions can stay stable while reference material is updated more often. Separate breakpoints let requests reuse the longest eligible prefix that remains unchanged.
234 293
235Each request can create up to four new cache writes. Breakpoints from earlier conversation turns are read-only: they can match the cache, but the request does not write them again. If more than four breakpoints are set, only the last four are written.294### Choose a caching mode
236 295
237In `implicit` mode, the breakpoint on the latest message uses one write slot. Up to the latest three explicit breakpoints can use the remaining slots. In `explicit` mode, up to the latest four explicit breakpoints can create new cache writes.
238 296
239For cache reads, OpenAI considers up to the latest 50 breakpoints in the conversation. When several breakpoints match cached content, the service reads from the longest matching prefix.
240 297
241### Improve cache matching with a prompt cache key298On GPT-5.6 and later, two controls determine where cache breakpoints are placed: `prompt_cache_options.mode` selects implicit or explicit-only caching, and `prompt_cache_breakpoint` marks a boundary you choose.
242 299
243Set `prompt_cache_key` on requests that share long, common prompt prefixes. Reuse the same key for those requests to help route them to the same cache and improve cache hit rates. Common values for `prompt_cache_key` include session IDs and user IDs.300- **Place breakpoints automatically.** Use implicit caching to place a breakpoint at the end of the latest eligible message. This is convenient for multi-turn threads that append to existing context.
301- **Choose breakpoints deliberately.** Place explicit markers at the end of stable content. Use explicit-only mode to avoid unnecessary cache writes for changing suffixes.
244 302
245For GPT-5.6, you must set `prompt_cache_key` to use the more reliable matching for both implicit and explicit caching. At each breakpoint, the service matches the key with the exact prompt prefix. Without a key, requests may still receive automatic cache hits, but they do not use the improved matching.
246 303
247Keep the total traffic across all prefixes for each key to approximately 15 requests per minute. If a key receives a higher rate, some requests may miss the cache. For higher-volume workloads, partition traffic across more keys and use a stable mapping so requests with the same key continue to share prefixes.
248 304
249### Measure cache reads and writes305> Illustration: In explicit-only mode, tools and schemas precede a stable developer-message prefix and breakpoint 1. One branch adds a variable developer suffix and more conversation turns before breakpoint 2, then splits into new user inputs. Another branch has an unselected variable suffix. Content after each branch's last selected breakpoint is charged at the uncached input rate without a cache-write charge.
250 306
251Monitor `cached_tokens` and `cache_write_tokens` to understand whether your breakpoint placement produces cache reuse or repeated cache writes.
252 307
253`cached_tokens` is the number of input tokens read from the cache. `cache_write_tokens` is the number of input tokens newly written to the cache.
254 308
255For the Responses API, both fields appear in `usage.input_tokens_details`. For the Chat Completions API, they appear in `usage.prompt_tokens_details`.
256 309
257```json310
258{311
259 "usage": {312
260 "input_tokens": 2600,313<a id="prompt-cache-key-best-practices"></a>
261 "input_tokens_details": {314
262 "cached_tokens": 2000,315
263 "cache_write_tokens": 400316
264 }317### Tune prompt cache keys
265 }318
266}319
320
321- **Group related requests.** Combine a prompt version with a stable user, workspace, session, or thread ID that matches how your application reuses context. For example:
322 - `prompt_name_v1:user_123` groups a user's related requests that share a prompt version.
323 - `prompt_name_v1:session_456` groups requests within one session.
324 - `prompt_name_v1:workspace_acme:shard_3` groups requests within a stable shard of a workspace.
325- **Keep keys stable.** Reuse the key while its prefix remains useful; do not generate a new key for every request.
326- **Split busy groups.** If a group receives high traffic and cache read hits decline, distribute it across more keys with a stable, deterministic mapping. Keep related requests on the same shard so they can reuse its cache.
327
328Create stable cache keys
329
330```javascript
331import { createHash } from "node:crypto";
332
333const tenantId = "acme";
334const sessionId = "session-42";
335const promptVersion = "support-v3";
336// Tune for peak traffic per tenant and reusable prompt group; monitor cache hits.
337const shardCount = 16;
338
339const digest = createHash("sha256")
340 .update(`${tenantId}:${sessionId}`)
341 .digest("hex");
342const shard = Number.parseInt(digest.slice(0, 8), 16) % shardCount;
343const promptCacheKey = `${promptVersion}:${tenantId}:shard-${shard}`;
267```344```
268 345
269In this example, 2,000 tokens were read from the cache and 400 additional tokens were written. The remaining 200 input tokens were neither read nor written. A longer cache write does not bill the already cached 2,000 tokens again.346```python
347import hashlib
270 348
271**Understand cache-write pricing**349tenant_id = "acme"
350session_id = "session-42"
351prompt_version = "support-v3"
352# Tune for peak traffic per tenant and reusable prompt group; monitor cache hits.
353shard_count = 16
354
355digest = hashlib.sha256(f"{tenant_id}:{session_id}".encode()).hexdigest()
356shard = int(digest[:8], 16) % shard_count
357prompt_cache_key = f"{prompt_version}:{tenant_id}:shard-{shard}"
358```
359
360```java
361import java.nio.charset.StandardCharsets;
362import java.security.MessageDigest;
363import java.util.HexFormat;
364
365String tenantId = "acme";
366String sessionId = "session-42";
367String promptVersion = "support-v3";
368int shardCount = 16;
369
370String digest =
371 HexFormat.of()
372 .formatHex(
373 MessageDigest.getInstance("SHA-256")
374 .digest((tenantId + ":" + sessionId).getBytes(StandardCharsets.UTF_8)));
375long shard = Long.parseLong(digest.substring(0, 8), 16) % shardCount;
376String promptCacheKey = promptVersion + ":" + tenantId + ":shard-" + shard;
377```
272 378
273Cache reads, cache writes, and ordinary input tokens are separate billing categories.
274 379
2751. Cached input tokens are billed at 0.1× the uncached input token rate.
2762. Tokens written to the cache are billed at 1.25× the uncached input token rate.
2773. Tokens that are neither read nor written are billed at the uncached input token rate.
278 380
279The 1.25× cache-write rate is the total rate for written tokens. It is not an additional charge on top of another full input-token charge. A breakpoint does not create a charge by itself. Charges apply to tokens that are actually written to the cache.
280 381
281Repeated writes increase cost when the resulting cache entries are not reused. If `cache_write_tokens` stays high while `cached_tokens` remains low, check whether an implicit breakpoint includes content that changes between requests.
282 382
283### Set cache lifetime
284 383
285Use `prompt_cache_options.ttl` to set the lifetime of all breakpoints written by a request. The only supported value is `30m`, which is also the default.384<a id="choose-a-cache-lifetime"></a>
286 385
287The 30-minute lifetime begins when the prefix is written and refreshes whenever the prefix is reused. A cached prefix remains eligible for reuse for 30 minutes after its most recent write or reuse, though OpenAI may retain it longer.
288 386
289Reusing a cached prefix refreshes its lifetime without creating another cache-write charge.
290 387
291### Troubleshoot common caching issues388### Configure cache retention
292 389
293- **`cached_tokens` is zero:** Check that the rendered prefix through the breakpoint contains at least 1,024 tokens. Confirm that an earlier request wrote the same prefix and that related requests use the same `prompt_cache_key`.
294 390
295- **Cache writes repeat on every request:** Check whether a timestamp, changing user input, tool-call history, or other request-specific content appears before the eligible breakpoint. Move the explicit breakpoint to the end of the stable prefix.
296 391
297- **Cache reads and writes are both nonzero:** In implicit mode, a request can read an earlier cached prefix and write newly appended content at the latest breakpoint. Use explicit-only mode if that new content should not be cached.392For earlier models, prefer setting `prompt_cache_retention` to `"24h"` for extended retention when the model and your data-retention requirements allow it. See [Cache lifetime](#cache-lifetime) for supported settings and defaults.
298 393
299- **Explicit mode produces no cache hits:** Confirm that at least one supported content block has `prompt_cache_breakpoint: { "mode": "explicit" }` and that the rendered prefix through the marker meets the 1,024-token minimum.
300 394
301- **Cache hits decrease at higher request volumes:** Keep traffic for each `prompt_cache_key` to approximately 15 requests per minute. Use stable, deterministic keys to partition larger workloads.
302 395
303- **A previously cached prompt no longer matches:** Check whether tool definitions, tool ordering, structured output schemas, images, prompt content, or request settings changed before the breakpoint.
304 396
305- **A breakpoint is rejected:** Attach the marker to a supported content block and use `explicit` as its mode. Do not attach a breakpoint to top-level Responses API `instructions`.
306 397
307## Prompt caching for earlier models398<a id="a-shared-prefix-just-below-the-caching-minimum"></a>
308 399
309Earlier models use automatic prompt caching to reuse matching prompt prefixes. When an eligible request is routed to a machine that recently processed the same prefix, the service can reuse the cached result instead of processing that content again.
310 400
311Prompt caching works automatically for supported models. A consistent `prompt_cache_key`, stable prompt structure, and an appropriate `prompt_cache_retention` setting can improve cache reuse.
312 401
313### How automatic prompt caching works402### Escape the minimum cacheable length cost trap
314 403
315Cache hits are only possible for exact prefix matches within a prompt. When a request arrives, the service checks whether an eligible initial portion of the prompt already exists in the cache on the selected machine.
316 404
317If a matching prefix is available, the service can reuse an eligible matching prefix and report those tokens in `cached_tokens`. If no match is available, the service processes the full prompt and may cache eligible content for future requests.
318 405
319Cache reuse is best-effort. A cache hit depends on the prompt prefix remaining identical, the cached content still being available, and the request reaching a machine that holds the matching entry.406If many requests reuse the same developer instructions and tool definitions, but that shared prefix falls below the model's [minimum cacheable length](#summary-of-model-differences), consider shortening it or expanding it with useful, stable instructions, examples, or reference material. Measure whether cache reuse offsets the additional input tokens and any cache-write charges, and ensure evaluations and behaviour remain stable.
320 407
321For example, separate requests can reuse shared instructions and reference material while the user message changes:408The chart highlights the minimum cacheable length cost trap where short prefix lengths can cost more uncached than expanding to the minimum cacheable token length.
322 409
323```text410#### Mathematical details
324Request 1: Shared instructions → Shared reference material → User message 1411
325Request 2: Shared instructions → Shared reference material → User message 2412
413
414For a cost-only comparison, let $$M$$ be the minimum cacheable length, $$L < M$$ the original prefix length, $$r$$ the cache-read multiplier, $$w$$ the cache-write multiplier, and $$N$$ the total number of requests. Assume the expanded prefix is exactly $$M$$ tokens, is written once, and is fully reused on every later request. In uncached-input-token equivalents, keeping the original prefix costs $$N \times L$$, while expanding it costs $$M \left[w + (N - 1)r\right]$$. The break-even original length is:
415
416$$
417L_{\mathrm{break\text{-}even}} = M\left(r + \frac{w-r}{N}\right)
418$$
419
420Expand when $$L > L_{\mathrm{break\text{-}even}}$$; keeping the shorter prefix costs less when $$L < L_{\mathrm{break\text{-}even}}$$. At equality, the costs are the same. The smallest whole-token length for which expansion is cheaper is $$\left\lfloor L_{\mathrm{break\text{-}even}} \right\rfloor + 1$$. Conversely, shrinking a cacheable prefix below $$M$$ loses caching: under the same assumptions, the shorter uncached prefix must fall below $$L_{\mathrm{break\text{-}even}}$$ to cost less than caching $$M$$ tokens. There is no universal maximum-cost prompt length; the crossover depends on reuse and pricing.
421
422For example, with $$M = 1{,}024$$, $$r = 0.1$$, and $$w = 1.25$$, the crossover is $$102.4 + \frac{1{,}177.6}{N}$$ tokens. Across 10 requests, expanding an original prefix of at least 221 tokens to 1,024 tokens is cheaper. As reuse grows, the crossover approaches 102.4 tokens. A 103-token prefix needs at least 1,963 total requests to benefit; a prefix of 102 tokens or fewer never does under these assumptions. This comparison excludes performance, output tokens, and unchanged request costs. Additional misses, writes, or different model rates change the result.
423
424
425
426
427
428
429
430
431
432<a id="monitor-cache-performance"></a>
433
434
435
436### Monitor cache performance
437
438
439
440- **Measure actual cache performance.** Track `usage.input_tokens_details.cached_tokens`, `usage.input_tokens_details.cache_write_tokens`, input-token counts, latency, and realized cost. Track the token cache-hit rate by dividing total cached tokens by total input tokens, aggregating both counts by user, workspace, day, or another useful grouping.
441- **Calculate input cost.** Use the token counts in `response.usage` and the model's [prices per million tokens](https://developers.openai.com/api/docs/pricing).
442- **Use the prompt caching dashboard.** Monitor cache hit rates in the [Prompt Caching Dashboard](https://platform.openai.com/usage?usage_section=prompt-caching).
443
444Calculate input cost
445
446```javascript
447function calculateInputCost(
448 usage,
449 inputPricePerMillion,
450 cacheInputMultiplier = 0.1,
451 cacheWriteMultiplier = 1.25
452) {
453 const inputTokens = usage.input_tokens;
454 const cachedTokens = usage.input_tokens_details.cached_tokens;
455 const cacheWriteTokens = usage.input_tokens_details.cache_write_tokens;
456 const ordinaryInputTokens = inputTokens - cachedTokens - cacheWriteTokens;
457
458 const weightedInputTokens =
459 ordinaryInputTokens +
460 cachedTokens * cacheInputMultiplier +
461 cacheWriteTokens * cacheWriteMultiplier;
462 const inputCost = (weightedInputTokens * inputPricePerMillion) / 1_000_000;
463 return inputCost;
464}
326```465```
327 466
328When the shared prefix is eligible and available, the second request can reuse that content without requiring additional request-specific cache configuration.467```python
468from openai.types.responses import ResponseUsage
469
470
471def calculate_input_cost(
472 usage: ResponseUsage,
473 input_price_per_million: float,
474 cache_input_multiplier: float = 0.1,
475 cache_write_multiplier: float = 1.25,
476) -> float:
477 input_tokens = usage.input_tokens
478 cached_tokens = usage.input_tokens_details.cached_tokens
479 cache_write_tokens = usage.input_tokens_details.cache_write_tokens
480 ordinary_input_tokens = input_tokens - cached_tokens - cache_write_tokens
481
482 weighted_input_tokens = (
483 ordinary_input_tokens
484 + cached_tokens * cache_input_multiplier
485 + cache_write_tokens * cache_write_multiplier
486 )
487 input_cost = weighted_input_tokens * input_price_per_million / 1_000_000
488 return input_cost
489```
329 490
330**Minimum cacheable prefix**
331 491
332The minimum cacheable prefix length varies by model and can range from 1,024 to 2,048 tokens. Prompts just above 1,024 tokens may not be cached consistently.
333 492
334Cache hits occur in increments of 128 tokens. The number of cached tokens can therefore be smaller than the full length of the shared prompt content.
335 493
336Make sure the repeated portion of the prompt meets the minimum for the model. A request can exceed the minimum overall and still fail to produce a cache hit if its matching prefix is too short.
337 494
338### Structure prompts for reuse
339 495
340Cache hits are only possible for exact prefix matches within a prompt. To realize caching benefits, place static content like instructions and examples at the beginning of your prompt, and put variable content, such as user-specific information, at the end. This also applies to images and tools, which must be identical between requests.
341 496
342Keep system or developer instructions, shared reference material, examples, tool definitions, and structured output schemas stable. Put user input, request identifiers, timestamps, and other changing content after the reusable prefix.
343 497
344If a dynamic value is needed only for logging or debugging, consider placing it in request metadata instead of inserting it into the prompt.498### Migrate prompt caching from an earlier model to GPT-5.6 and later
345 499
346**Keep tools and schemas identical**
347 500
348Tool definitions, tool ordering, and structured output schemas contribute to the prompt prefix. Changes to tool descriptions, parameter schemas, schema keys, or ordering can reduce cache reuse.
349 501
350When you need to restrict which tools are available on a particular request, keep the underlying `tools` array unchanged and use `allowed_tools` where supported.502- Keep existing stable prefixes.
503- Keep existing `prompt_cache_key` values.
504- Replace `prompt_cache_retention` with `prompt_cache_options.ttl`.
505- Confirm that reusable prefixes meet the model's [minimum cacheable length](#summary-of-model-differences).
506- If the default breakpoint includes content that changes between requests, add an explicit breakpoint after the stable prefix.
507- Use `prompt_cache_options.mode: "explicit"` when later content is not worth writing.
508- Compare `cached_tokens`, `cache_write_tokens`, latency, and total cost before and after migration.
351 509
352**Preserve conversation history**
353 510
354For multi-turn conversations, append new user and assistant messages instead of rewriting earlier messages. Changing, deleting, or reordering earlier content changes the prefix and can cause a cache miss.
355 511
356Context truncation, summarization, and compaction can reduce prompt size, but they can also reset the reusable prefix. Balance the savings from shorter prompts against the loss of existing cache reuse.
357 512
358### Improve cache hit rates with a prompt cache key
359 513
360Set `prompt_cache_key` on requests that share long, common prompt prefixes. Reuse the same key for those requests to help route them to the same cache and improve cache hit rates.514## Examples
361 515
362Requests are routed based on the initial prompt prefix. When you provide `prompt_cache_key`, it is combined with the prefix hash, allowing you to influence routing. This is especially beneficial when many requests share long, common prefixes.516<a id="single-turn-llm-as-a-judge"></a>
363 517
364Keep the total traffic across all prefixes for each key to approximately 15 requests per minute. If a key receives a higher rate, some requests may miss the cache. For higher-volume workloads, partition traffic across more keys and use a stable mapping so requests with the same key continue to share prefixes.
365 518
366A cache key improves routing but does not make different prompt prefixes match. Keep the prefix and the cache key consistent across requests that should share cached content.
367 519
368<a id="prompt-cache-retention"></a>520### Single-turn LLM-as-a-Judge
369 521
370### Configure prompt cache retention
371 522
372Use `prompt_cache_retention` to select the retention policy for a supported Responses API or Chat Completions request. Available values depend on the model.
373 523
374For models that support both in-memory and extended retention, prompt cache pricing is the same for both policies.524Consider a single-turn LLM judge that determines whether a completed interaction shows evidence that the user is satisfied after an interaction with a chatbot. Each request uses the same grading rubric and labeled few-shot examples to evaluate a different interaction.
375 525
376**In-memory prompt cache retention**526- **Preserving the prefix:** The fixed rubric and examples come first. Their combined length is deliberately kept just above the model's [minimum cacheable length](#summary-of-model-differences), using material that helps calibrate the judge. The interaction being evaluated comes last.
527- **Prompt cache key:** A stable `prompt_cache_key`, such as `satisfaction_judge_v1`, groups requests using the same rubric version.
528- **Caching mode and breakpoint:** Explicit-only caching is enabled, with a breakpoint after the fixed rubric and examples. The user–chatbot conversation being evaluated comes after that breakpoint and is not written to the cache, avoiding a cache-write charge for content that is unlikely to be reused.
377 529
378In-memory prompt cache retention is available for models that accept `prompt_cache_retention: "in_memory"`.530For illustration, a deployment using these principles might achieve a **token cache-hit rate of around 70%**. This is a hypothetical figure, not a measured deployment result. Actual cache-hit rates depend on your context and application usage.
379 531
380When using the in-memory policy, cached prefixes generally remain active for 5 to 10 minutes of inactivity, up to a maximum of one hour. In-memory cached prefixes are held only in volatile memory.532Responses API request for a single-turn judge
381 533
382<a id="extended-prompt-cache-retention"></a>534```json
535{
536 "model": "gpt-5.6-sol",
537 "reasoning": { "effort": "medium", "context": "all_turns" },
538 "text": { "verbosity": "low" },
539 "prompt_cache_key": "satisfaction_judge_v1",
540 "prompt_cache_options": { "mode": "explicit" },
541 "input": [
542 {
543 "role": "developer",
544 "content": [
545 {
546 "type": "input_text",
547 "text": "Judge whether the completed interaction provides evidence that the user is satisfied. Return true or false. Full grading rubric and labeled few-shot examples...",
548 "prompt_cache_breakpoint": { "mode": "explicit" }
549 }
550 ]
551 },
552 {
553 "role": "user",
554 "content": "Completed interaction to evaluate..."
555 }
556 ]
557}
558```
383 559
384**Extended prompt cache retention**
385 560
386Extended prompt cache retention keeps cached prefixes active for longer, up to a maximum of 24 hours.
387 561
388The 24-hour period is a maximum, not a guarantee that every request will receive a cache hit. Reuse still depends on an exact matching prefix, cache availability, and request routing.
389 562
390**Models that support extended retention**
391 563
392Extended prompt cache retention is available for the following models:
393 564
394- `gpt-5.5`565<a id="customer-support-agent"></a>
395- `gpt-5.5-pro`
396- `gpt-5.4`
397- `gpt-5.2`
398- `gpt-5.1-codex-max`
399- `gpt-5.1`
400- `gpt-5.1-codex`
401- `gpt-5.1-codex-mini`
402- `gpt-5.1-chat-latest`
403- `gpt-5`
404- `gpt-5-codex`
405- `gpt-4.1`
406 566
407**Retention defaults and Zero Data Retention**
408 567
409For `gpt-5.5` and `gpt-5.5-pro`, only `24h` is supported through `prompt_cache_retention`.
410 568
411For models that support both `in_memory` and `24h`, the default depends on your organization's data retention policy:569### Multi-turn agent
412 570
413- Organizations without Zero Data Retention enabled default to `24h`.
414- Organizations with Zero Data Retention enabled default to `in_memory` when `prompt_cache_retention` is not specified.
415 571
416Verify the available retention policies for your model and organization before selecting a value.
417 572
418### Measure cache hits and costs573Consider a multi-turn agent with long, shared developer instructions and frequent tool calls. Typical usage sees users running multiple sessions with the agent at once, and often forking the threads.
419 574
420Use `cached_tokens` to see how many input tokens were read from the cache. The field is present even when no tokens were cached.575- **Preserving the prefix**: Each turn appends new messages, tool calls, and results without rewriting earlier context, so the reusable prefix grows over time.
576- **Prompt cache key:** The `prompt_cache_key` is defined for each user-agent pair, shared across that user's sessions with the agent. For example, `agent_123_v1:user_456` groups user 456's sessions and forks with agent 123. The session and thread IDs are kept out of the key when those sessions should share the same reusable prefix.
577- **Implicit caching mode:** Implicit caching is enabled so the latest eligible user or tool message provides a breakpoint.
578- **Explicit breakpoints:** A breakpoint is added after each tool result to preserve earlier reusable prefixes and improve cache efficiency of forking.
421 579
422For the Responses API, the field appears in `usage.input_tokens_details.cached_tokens`. For the Chat Completions API, it appears in `usage.prompt_tokens_details.cached_tokens`.580An example deployment using these principles reported a **token cache-hit rate >90%**. This figure illustrates a possible outcome. Actual cache-hit rate ceilings will depend upon your own context and application usage.
423 581
424The following Chat Completions usage example shows a request that reused 1,920 of its 2,006 prompt tokens:582Responses API request for a multi-turn agent
425 583
426```json584```json
427{585{
428 "usage": {586 "model": "gpt-5.6-sol",
429 "prompt_tokens": 2006,587 "reasoning": { "effort": "medium", "context": "all_turns" },
430 "completion_tokens": 300,588 "text": { "verbosity": "medium" },
431 "total_tokens": 2306,589 "prompt_cache_key": "agent_123_v1:user_456",
432 "prompt_tokens_details": {590 "prompt_cache_options": { "mode": "implicit" },
433 "cached_tokens": 1920591 "tools": [
592 {
593 "type": "function",
594 "name": "function_name",
595 "description": "Function description",
596 "parameters": { "...": "..." }
434 }597 }
598 ],
599 "input": [
600 {
601 "role": "developer",
602 "content": "Stable developer instructions and reference material..."
603 },
604 { "role": "user", "content": "Can you do...?" },
605 {
606 "type": "function_call",
607 "call_id": "call_123",
608 "name": "function_name",
609 "arguments": "..."
610 },
611 {
612 "type": "function_call_output",
613 "call_id": "call_123",
614 "output": [
615 {
616 "type": "input_text",
617 "text": "Tool result...",
618 "prompt_cache_breakpoint": { "mode": "explicit" }
435 }619 }
620 ]
621 },
622 { "role": "assistant", "content": "Assistant response..." },
623 { "role": "user", "content": "Can you also do...?" }
624 ]
436}625}
437```626```
438 627
439In this example, the remaining 86 prompt tokens were not read from the cache. Monitor cached-token usage across requests to identify changes in prompt structure, traffic patterns, or cache availability.
440 628
441**Pricing and rate limits**
442 629
443Creating a cache entry has no additional fee. Cached input is billed at the cached-input rate when the model offers one. Rates and discounts vary by model.
444 630
445Cached input tokens still count toward tokens-per-minute rate limits. Prompt caching does not change rate-limit calculations or guarantee identical model outputs.
446 631
447### What can be cached
448 632
449- **Messages:** System, developer, user, and assistant messages can contribute to a reusable prompt prefix.633<a id="troubleshooting"></a>
450- **Images:** Image inputs can be cached when the images, their order, and their detail settings remain the same.634
451- **Tools:** Tool definitions, descriptions, parameter schemas, and tool ordering can contribute to the prefix.635## Gotchas
452- **Structured outputs:** A structured output schema can be included in the reusable prompt prefix.636
453- **Audio:** Supported audio inputs can contribute to cacheable prompt content.637
638
639### A shared prefix is not always a cached prefix
640
641
642
643This is particularly prevalent when migrating from earlier models to GPT-5.6 or later due to the change in implicit caching behaviour. If requests share a long prefix but have different suffixes, caching the first complete request implicitly-only does not make the shorter shared prefix reusable.
644
645Consider a static developer message followed by a dynamic user message in each request. This request writes through the dynamic content. Changing that content in the next request does not match the longer cached prefix, and there is no separate breakpoint after the static content.
646
647Without a breakpoint after the static content
648
649```json
650{
651 "model": "gpt-5.6-sol",
652 "reasoning": { "effort": "medium", "context": "all_turns" },
653 "text": { "verbosity": "low" },
654 "prompt_cache_key": "prompt_name_v1",
655 "prompt_cache_options": { "mode": "implicit" },
656 "input": [
657 { "role": "developer", "content": "Static content..." },
658 { "role": "user", "content": "Dynamic content..." }
659 ]
660}
661```
662
663
664To remediate, place an explicit breakpoint after the static content in both requests. The first request writes the reusable prefix; the next can reuse it even when the dynamic content changes. This example uses explicit-only mode to avoid writing the dynamic content to cache.
665
666With a breakpoint after the static content
667
668```json
669{
670 "model": "gpt-5.6-sol",
671 "reasoning": { "effort": "medium", "context": "all_turns" },
672 "text": { "verbosity": "low" },
673 "prompt_cache_key": "prompt_name_v1",
674 "prompt_cache_options": { "mode": "explicit" },
675 "input": [
676 {
677 "role": "developer",
678 "content": [{
679 "type": "input_text",
680 "text": "Static content...",
681 "prompt_cache_breakpoint": { "mode": "explicit" }
682 }]
683 },
684 { "role": "user", "content": "Dynamic content..." }
685 ]
686}
687```
688
689
690
691
692
693
694
695
696### Minimum cacheable length varies by model
697
698
699
700A prefix that qualifies for caching on one model may be too short on another. Check the [model comparison](#summary-of-model-differences) and measure the reusable prefix with the model and settings you actually use. When changing models, repeat that check rather than assuming the previous model's threshold still applies.
701
702
703
704
705
706
707
708### Compaction can reduce cache reuse
709
710
711
712[Compaction](https://developers.openai.com/api/docs/guides/compaction) replaces earlier conversation context with a shorter representation. That can change the prefix, so the first request after compaction may reuse less of the previous cache even when the conversation is logically the same.
713
714Keep reusable instructions and reference material stable where possible, then let subsequent turns build on the compacted context. Compare total input cost before and after compaction: fewer input tokens can still save money even when the cache-hit rate falls.
715
716
717
454 718
455All reusable content must remain identical across requests. Changes earlier in the prompt can invalidate reuse for the content that follows.
456 719
457## Frequently asked questions720## Frequently asked questions
458 721
4591. **How is data privacy maintained for caches?**
460 722
461 Prompt caches are not shared between organizations. Only members of the same organization can access caches of identical prompts. Cache data handling depends on the model and retention policy. See the [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for the current application-state, Zero Data Retention, and data residency details.
462 723
4632. **Does Prompt Caching affect output token generation or the final response of the API?**724### Does prompt caching affect output generation?
725
726
727
728No. Prompt caching does not change how the model generates output tokens. The model generates a new response using the cached prefix, so identical requests are not guaranteed to produce identical outputs.
729
730
731
732
733
734
735
736### Can I manually clear the cache?
737
738
739
740No. Manual cache clearing is not currently available. Cache entries expire according to the model's [cache lifetime](#cache-lifetime) and retention settings.
741
742
743
464 744
465 Prompt Caching does not change how the model generates output tokens. The model computes a new response from the cached prompt prefix, so otherwise identical nondeterministic requests are not guaranteed to return identical output.
466 745
4673. **Is there a way to manually clear the cache?**
468 746
469 Manual cache clearing is not currently available. For models before the GPT-5.6 family that use in-memory retention, typical cache evictions occur after 5-10 minutes of inactivity, though entries can remain for up to one hour during off-peak periods. For GPT-5.6 models and later model families, cached prefixes remain eligible for reuse for 30 minutes and may be retained longer.
470 747
4714. **Will I be expected to pay extra for writing to Prompt Caching?**748### Do cached prompts count toward rate limits?
472 749
473 Cache writes have no additional fee on models before the GPT-5.6 family. On GPT-5.6 models and later model families, cache writes are billed at 1.25× the uncached input token rate and reported in `cache_write_tokens`. Cache reads continue to be reported in `cached_tokens`.
474 750
4755. **Do cached prompts contribute to TPM rate limits?**
476 751
477 Yes, as caching does not affect rate limits.752Yes. Cached input tokens still count toward tokens-per-minute limits. Prompt caching does not change how [rate limits](https://developers.openai.com/api/docs/guides/rate-limits) are calculated.