OpenAI Prompt Cache Diagnostics Tutorial: Fix Cache Misses
Learn how to use OpenAI Prompt Cache Diagnostics in the Responses API to find cache misses, fix unstable prefixes, and measure real cache savings.
On this page
A long prompt can look identical in your application while OpenAI quietly stops reusing its cached prefix. With GPT-5.6 and later supported models, the Responses API now gives developers a way to investigate that problem instead of guessing: Prompt Cache Diagnostics compares a current request with an earlier response and reports what changed.
This tutorial shows how to use that diagnostic flow, read the result, and fix the request structure that caused the miss. The important distinction is that diagnostics do not make caching happen; they explain why an expected reusable prefix was not reused.
Why prompt cache misses are difficult to find
Prompt caching works by reusing an identical prefix from an earlier request. That prefix can contain developer instructions, reference material, conversation history, tool definitions, and other input that appears before the changing part of the request. For GPT-5.6 and later, a cacheable prefix must contain at least 1,024 visible input tokens, and the relevant request settings must remain compatible.
The awkward part is that a cache miss does not tell you which part changed. A developer may add a tool, reorder two tools, modify an output schema, change reasoning effort, edit an earlier message, or accidentally insert a timestamp into otherwise stable instructions. The resulting request still looks almost the same to a human reviewing application code, but the prefix no longer matches exactly where it needs to.
That matters especially in agents. A multi-turn agent can repeatedly send the same instructions and tool definitions while adding a small amount of new user input on every turn. If the application rewrites old context instead of appending new information, the reusable prefix can disappear. The same cost problem can show up in larger agent calls, where repeated context is a substantial part of the input.
What OpenAI Prompt Cache Diagnostics actually compares
Prompt Cache Diagnostics uses a completed response as a baseline. You pass that response's ID through prompt_cache_options.comparison_response_id on the next request. OpenAI then compares the current request with the earlier request and returns a diagnostic object when it can identify a relevant difference.
The comparison does not load the earlier conversation into the new request, and it does not change how the current request is processed. It is an observability feature. Your application still performs normal prompt caching independently of the diagnostic comparison.
There are several possible outcomes. A cache_hit means the comparison found no cache-affecting difference. A cache_miss means a difference was found and includes a reason. comparison_response_not_found means the baseline diagnostic record is unavailable, while unavailable means the comparison could not produce a conclusive result.
Set up a baseline Responses API request
Start with a request containing enough reusable material to create a meaningful cacheable prefix. In a real application this might be a support policy, product documentation, agent instructions, or a large collection of stable reference material. The following Python example deliberately keeps the structure simple so the second request has something useful to compare.
from pathlib import Path
from openai import OpenAI
client = OpenAI()
policy = Path("support-policy.txt").read_text()
first = client.responses.create(
model="gpt-6-astra",
instructions=policy,
input="Reply with exactly OK.",
tools=[
{
"type": "function",
"name": "get_time",
"description": "Get the current time.",
"parameters": {
"type": "object",
"properties": {}
}
}
],
)
print(first.id)Save the returned response ID. In production, you would normally keep it with the request telemetry that lets you associate one completed turn with the next request. The baseline should be recent enough for the diagnostic record to remain available, and it must belong to the same organization as the request being compared.
The reusable prefix also needs to meet the model's cache requirements. For GPT-5.6 and later supported models, the documented minimum is 1,024 visible input tokens. A tiny test prompt is therefore useful for checking API syntax, but it is not a useful demonstration of a real cacheable workload.
Compare the next request with the baseline
Now make a second request that is intentionally almost identical. Change one cache-sensitive detail so the diagnostic has something to find. Here the function name changes from get_time to get_date.
second = client.responses.create(
model="gpt-6-astra",
instructions=policy,
input="Reply with exactly OK.",
tools=[
{
"type": "function",
"name": "get_date",
"description": "Get the current date.",
"parameters": {
"type": "object",
"properties": {}
}
}
],
prompt_cache_options={
"comparison_response_id": first.id
},
)
diagnostics = second.prompt_cache_diagnostics
if diagnostics is not None:
print(diagnostics.type)
```
if diagnostics.type == "cache_miss":
print(diagnostics.reason)
print(diagnostics.comparison_reusable_tokens)
print(diagnostics.cache_missed_tokens)The useful field is reason. In this example, changing the tool definition can produce a tools_changed diagnosis. The response can also report how many reusable tokens were available in the comparison and how many of those tokens were missed.
Do not confuse those diagnostic counts with the billing counters from the current request. When measuring actual cache reuse, inspect usage.input_tokens_details.cached_tokens. Diagnostics explain the difference between requests; usage fields tell you what the current request actually reused.
Read the cache miss reason before changing your code
Once a miss is identified, resist the temptation to change several things at once. The diagnostic reason tells you where to start. If the model changed, check routing, fallback logic, and experiments. If the service tier changed, make sure requests intended to share a prefix use compatible tiers. If the tools changed, compare names, descriptions, schemas, ordering, and tool-specific configuration.
Other reasons can point to changes in structured output, reasoning effort, response verbosity, conversation compaction, or the input itself. An input change can be surprisingly small: a timestamp, request identifier, edited earlier message, reordered history, or removed message can alter the prefix from that point onward.
| Diagnostic reason | What to inspect |
|---|---|
model_changed | Model selection, routing, experiments, and fallbacks |
tools_changed | Tool names, descriptions, schemas, ordering, and configuration |
text_format_changed | Structured output format and schema |
reasoning_effort_changed | Reasoning configuration between requests |
verbosity_changed | Output verbosity configuration |
context_compacted | Whether earlier conversation content was replaced |
input_changed | Instructions, history, tool results, timestamps, or other prefix content |
The exact diagnostic is more useful than a generic rule such as “make your prompts consistent.” It gives you a concrete place to investigate, after which you can repeat the comparison to see whether the next problem has been exposed.
Fix the most common problem: unstable prompt prefixes
The safest structure is to put information that changes rarely before information that changes frequently. Stable instructions, examples, reference material, and fixed tool definitions should form the reusable prefix. User-specific values, timestamps, request IDs, and the current question should come later.
For example, avoid constructing one giant developer message that begins with stable instructions and then inserts the current time somewhere in the middle. Once that timestamp changes, everything after it is different too. Instead, keep the stable material together and place changing values after the reusable boundary.
{
```
"model": "gpt-5.6",
"prompt_cache_options": {
"mode": "explicit"
},
"input": [
{
"role": "developer",
"content": [
{
"type": "input_text",
"text": "Stable instructions and shared reference material...",
"prompt_cache_breakpoint": {
"mode": "explicit"
}
}
]
},
{
"role": "developer",
"content": "Dynamic instructions, timestamps, or user-specific information..."
},
{
"role": "user",
"content": "The current question..."
}
]
}This is where explicit cache breakpoints become useful. GPT-5.6 and later supported model families let you mark the end of a reusable prefix with prompt_cache_breakpoint. With mode set to explicit, only the breakpoints you choose are used for cache writes and reads, allowing changing suffixes to remain uncached when repeatedly writing them would not be useful.
Keep tools stable in an AI agent
Tool definitions are part of the request structure that can affect cache reuse. Changing a function's name, description, schema, or ordering can invalidate the reusable prefix. This is easy to overlook because application developers often generate tool arrays dynamically.
If your application has a fixed set of tools, construct them deterministically and keep their order stable. If you need to disable some tools for a particular turn without changing the supplied tool definitions, OpenAI's documentation recommends mechanisms such as tool choice or allowed tools rather than rebuilding the tool list.
This becomes especially relevant when an agent grows over time. A developer may add one new function to an existing tool registry and assume only the new capability has changed. From the cache's perspective, however, the request's tool portion has changed, so a previously reusable prefix may no longer match.
Use explicit breakpoints when the suffix changes often
Implicit caching is convenient because OpenAI can choose eligible breakpoint locations automatically. It works well for many multi-turn conversations where new messages are simply appended. Explicit caching is more useful when you know exactly which part of your request is stable and want to prevent frequently changing content from becoming part of a cache write.
GPT-5.6 and later support up to four cache writes per request. Multiple breakpoints can therefore represent sections that change at different rates. A large shared policy might be stable for weeks, tool results might change every turn, and another reference block might change once per session. Separating those boundaries gives the application more control than treating the entire prompt as one cacheable unit.
There is a cost trade-off here. For GPT-5.6 and later, cache writes are charged at 1.25 times the standard uncached input-token rate, while cache reads cost 0.1 times that rate. A cache write is therefore worth considering when the same prefix will be reused, not simply because the prefix happens to be large.
Measure whether the fix actually saved money
After changing the request structure, send representative traffic rather than relying on one successful test. Compare cached_tokens, cache_write_tokens, total input tokens, latency, and overall input cost across several requests. A higher cache-hit count is useful only when the workload actually benefits from the reused context.
For example, suppose a request contains a large stable knowledge base followed by a short user question. Moving the question after an explicit breakpoint can make the knowledge base reusable across requests. If users rarely repeat that same knowledge base, however, writing it to the cache may add cost without enough subsequent reads to compensate.
The practical target is not “maximum caching.” It is a prompt structure where expensive, stable context is reused while short-lived information stays out of the reusable prefix.
Diagnose multi-turn cache misses systematically
For a conversation or agent, keep the response ID from each completed turn and compare the next turn against the appropriate baseline. If the diagnostic reports a miss, fix the identified difference and run the comparison again. Continue until the request structure is stable enough for the intended prefix to be reusable.
Keep the original conversation history append-only where possible. Rewriting earlier messages, inserting content into an existing message, or compacting context can change the prefix. Compaction can still be worthwhile because fewer total input tokens may save more money than a lower cache-hit rate costs, so the correct measurement is total input cost rather than cache percentage alone.
A useful production workflow is therefore simple: record the request configuration, record actual cached-token usage, keep recent response IDs available for diagnosis, investigate unexpected drops, and verify the result after each structural change. This turns prompt caching from something the application merely hopes will work into something the engineering team can observe and tune.
What Prompt Cache Diagnostics cannot tell you
Diagnostics are deliberately narrower than a general request debugger. A cache_hit comparison does not mean every input token was cached; new tokens may still need processing. Likewise, an unavailable result is not evidence of either a hit or a miss. It means the diagnostic could not produce a conclusive comparison.
The diagnostic record can also expire, so an old response ID may return comparison_response_not_found. OpenAI says the diagnostic system uses configuration metadata, token-count estimates, and hashes rather than storing raw prompts or model outputs for this feature, and the feature is compatible with Zero Data Retention.
Most importantly, the diagnostic does not repair the request. If it reports input_changed, your application still has to decide whether the changed input was intentional or whether the changing content should have been moved after a cache breakpoint. That distinction belongs to the application architecture, not the diagnostic API.
The practical pattern to keep
For a production AI application, start by making the stable prefix genuinely stable. Keep shared instructions and reference material together, keep tool definitions deterministic, append conversation turns instead of rewriting them, and put volatile information after the reusable boundary. Use explicit breakpoints when the natural message boundaries do not match the way your context changes.
Then use Prompt Cache Diagnostics when the numbers disagree with your expectations. Compare the current response with a recent baseline, read the first reported reason, fix one structural difference, and measure actual cached tokens and cost again. That workflow is much more reliable than repeatedly rearranging prompts until cached_tokens happens to rise.
The useful shift is not simply cheaper prompts. It is being able to explain why a long-running AI application stopped reusing context and verify that the fix worked. Once cache behavior is treated as an observable part of the request architecture, prompt design becomes an engineering problem you can measure rather than a performance mystery.
Written by


