Why is my prompt cache hit rate low? Prefix breakers in OpenAI and Anthropic requests
Prompt caches reuse only an identical prefix. Check timestamps, tools, compaction, and TTL, and compute the hit rate correctly, including on AIVAX.
Common causes of a low prompt cache hit rate: something at the start of the request changes between calls, the tool list or model settings change, earlier turns are rewritten (summarized, compacted, truncated, edited), or the next request arrives after the cache entry expired. Caches reuse an exact prefix, so an early change prevents reuse of everything after it, although the unchanged part before it can still be reused. Before tuning anything, confirm that you are computing the rate correctly, because providers and gateways report cached tokens differently.
This guide covers the formulas, the usual prefix breakers on OpenAI and Anthropic, why compaction lowers the rate (and when that is fine), and which AIVAX gateway behaviors change the prefix. It does not give measured hit rates. Everything about OpenAI and Anthropic comes from their documentation as read on 9 October 2026; everything about AIVAX comes from the gateway code and the maintained docs.
How do you calculate the hit rate correctly?
Divide cached input tokens by total input tokens, summed over the requests you care about. The catch is what "total input" means in each response.
| API shape | Cached tokens | Total input tokens |
|---|---|---|
| OpenAI Responses API | usage.input_tokens_details.cached_tokens | usage.input_tokens (includes cached tokens) |
| Anthropic Messages API | usage.cache_read_input_tokens | cache_read_input_tokens + cache_creation_input_tokens + input_tokens. input_tokens counts only what comes after the last breakpoint |
AIVAX chat/completions | usage.prompt_tokens_details.cached_tokens | prompt_tokens + cached_tokens. The envelope subtracts cached tokens from prompt_tokens |
AIVAX uses a different denominator from OpenAI here. In the gateway code, the OpenAI-compatible response reports prompt_tokens as the prompt tokens minus the cached ones, and total_tokens still includes them. A response with prompt_tokens: 84, cached_tokens: 1792, and completion_tokens: 16 has total_tokens: 1892. Its hit rate is 1792 / (84 + 1792), about 96%. Dividing by prompt_tokens alone gives 2,133%, a sign that you are using the OpenAI formula on a different shape. This is an illustration of the field semantics, not a measurement.
Two more details apply on AIVAX. The gateway sums the provider usage of every model call it makes for one request, so a request that runs a tool loop reports the aggregate, not the last call. And when the model has a cached-input price, the gateway bills cached tokens at that rate and the remaining input tokens at the regular input rate (pricing lists model rates).
This script reads one usage object per line and prints the aggregate and per-request rates. Save it as hit-rate.mjs and run node hit-rate.mjs aivax usage.jsonl.
import { readFileSync } from "node:fs";
const [shape, file] = process.argv.slice(2);
const readers = {
"openai-responses": (u) => ({
cached: u.input_tokens_details?.cached_tokens ?? 0,
total: u.input_tokens,
}),
anthropic: (u) => {
const cached = u.cache_read_input_tokens ?? 0;
return {
cached,
total: cached + (u.cache_creation_input_tokens ?? 0) + u.input_tokens,
};
},
aivax: (u) => {
const cached = u.prompt_tokens_details?.cached_tokens ?? 0;
return { cached, total: u.prompt_tokens + cached };
},
};
const read = readers[shape];
if (!read || !file) {
console.error("usage: node hit-rate.mjs <openai-responses|anthropic|aivax> usage.jsonl");
process.exit(1);
}
const rows = readFileSync(file, "utf8")
.split("\n")
.filter(Boolean)
.map((line) => read(JSON.parse(line)));
const pct = (r) => ((100 * r.cached) / r.total).toFixed(0);
const cached = rows.reduce((n, r) => n + r.cached, 0);
const total = rows.reduce((n, r) => n + r.total, 0);
console.log(`${rows.length} requests: ${cached} of ${total} input tokens cached (${((100 * cached) / total).toFixed(1)}%)`);
console.log(`per request: ${rows.map((r) => `${pct(r)}%`).join(" ")}`);
Look at the per-request line, not only the aggregate. A healthy session that falls to 0% on one request points to a specific change; a flat low rate points to a design problem, such as a prefix below the minimum cacheable length (1,024 tokens on GPT-5.6 and later; 512 to 4,096 on Claude depending on the model).
What breaks the prefix?
Both providers render the request in a fixed order and match from the start. Anthropic documents that order as tools, then system, then messages, and a change at one level invalidates that level and everything after it. OpenAI describes the cached prefix as the full rendered context: hidden instructions, tool definitions, developer messages, and history.
| Change between requests | What the documentation says | Fix |
|---|---|---|
| Timestamp, request ID, or user name near the top | Both: content after the changed text cannot match. Anthropic's example shows a per-request timestamp causing a fresh cache write on every call and no reads | Move dynamic text after the stable instructions, or into a later message |
| Tool added, removed, reordered, or edited | OpenAI lists tool names, descriptions, schemas, and ordering as prefix-affecting. Anthropic: a tool change invalidates tools, system, and messages | Keep the list fixed. Use tool_choice: "none" or allowed_tools (OpenAI) instead of removing definitions |
tool_choice, images, thinking, or effort changed | Anthropic: each invalidates the message blocks, and for thinking and effort possibly tools and system, depending on the model | Keep these constant within a conversation |
| Reasoning effort or output schema changed | OpenAI: reasoning.effort, text.verbosity, and text.format can change the rendered instructions | On GPT-6 models, append a configuration_update item to change effort without rewriting the prefix |
| Model changed | OpenAI: a different model can use different weights and caching behavior | Pin the model within a conversation (see pin a model ID or use an alias) |
| Earlier turns summarized, compacted, truncated, or edited | OpenAI: these can change the prefix and reset reuse. Extending an earlier message instead of appending a new one can also prevent reuse | Append; rewrite history rarely and deliberately |
| Next request arrives after expiry | See the lifetime section below | Longer retention, or a request pattern that stays inside the window |
Anthropic adds one trap for long conversations: the cache lookup checks at most 20 blocks back from a breakpoint, and it finds only entries that earlier requests wrote at their own breakpoints. If a conversation grows by 20 or more blocks between requests, an automatic breakpoint can miss the last write. Anthropic's guidance is to place cache_control on the last block that stays identical across requests, and to add a second breakpoint closer to the end when the history grows fast.
How do you find the first request that changed?
For a coding agent or any long session, the useful question is which field differed between two consecutive requests. On GPT-5.6 and later in the Responses API, OpenAI can answer it: set prompt_cache_options.comparison_response_id to the earlier response's id, and the new response includes prompt_cache_diagnostics with a reason such as tools_changed, context_compacted, or input_changed. Diagnostics are best effort and cover only those models, so you also need a provider-independent check.
If you securely capture your outgoing request bodies, this script compares two of them. Treat those files as sensitive: they can contain private messages, tool results, or credentials embedded in prompts. The script reports changed request-level settings and the index of the first message that differs. Save it as first-break.mjs and run node first-break.mjs previous.json current.json.
import { readFileSync } from "node:fs";
const [previous, current] = process.argv.slice(2).map((f) => JSON.parse(readFileSync(f, "utf8")));
const same = (x, y) => JSON.stringify(x) === JSON.stringify(y);
const settings = [
"model", "tools", "tool_choice", "parallel_tool_calls", "response_format",
"text", "reasoning", "reasoning_effort", "verbosity", "thinking", "system",
];
for (const key of settings) {
if (!same(previous[key], current[key])) console.log(`differs before the conversation: ${key}`);
}
const before = previous.messages ?? previous.input ?? [];
const after = current.messages ?? current.input ?? [];
const shared = Math.min(before.length, after.length);
let i = 0;
while (i < shared && same(before[i], after[i])) i++;
console.log(
i === shared
? `first ${shared} items identical (${before.length} before, ${after.length} now)`
: `first changed item: index ${i} (role: ${after[i].role})`,
);
Index 0 with a system or developer role almost always means dynamic text in the instructions. A differs before the conversation line for tools means the list changed, and nothing after it could be reused. Identical items only rule out changes in the fields this script checks. It does not compare the Responses API instructions field or resolve server-side history referenced by previous_response_id; check those before investigating lifetime and minimum length.
Why does compaction drop the hit rate, and is that a problem?
Compaction replaces earlier turns with a shorter representation, so the prefix changes from the first replaced token onward. OpenAI says this directly in its prompt-caching guide: the first request after compaction may reuse less of the previous cache, and context_compacted is one of its diagnostic reasons. It also says to compare total input cost before and after, because fewer input tokens can still cost less even when the hit rate falls.
That is the right test, and the arithmetic is easy to run. Assume, for illustration only, that cached tokens cost 10% of the regular input price and ignore cache-write surcharges. A request with 150,000 input tokens at a 90% hit rate costs the equivalent of 150,000 × (0.9 × 0.1 + 0.1) = 28,500 regular tokens. After compaction, 20,000 tokens at a 40% hit rate cost 20,000 × (0.4 × 0.1 + 0.6) = 12,800. The hit rate fell from 90% to 40%, and the illustrative per-request input cost fell by 55%. That excludes the cost of producing the compacted context and any output charges. The real discount and write pricing differ by model, so use the provider's price list rather than these numbers.
What you can control is what stays stable around the compaction: keep the system prompt and tool definitions byte-identical, and let later turns build on the compacted context instead of compacting again on every request.
Which AIVAX gateway behaviors change the prefix?
An AIVAX AI Gateway builds the upstream request for you, so some prefix breakers come from gateway settings rather than from your client. These are read from the code and the maintained docs; I did not measure their effect on hit rates.
- Memory and reminders sit in the system message. When a request carries a user reference ID, the gateway appends saved memory items and upcoming reminders to the system instructions, after the instructions you configured. Saving a new memory changes that text, and the reminders shown depend on the current time. Both come before the conversation, so a change there changes the prefix for every message. See how agent memory works for the write path.
- Tool message truncation.
ToolContextCountkeeps only the last N tool responses and replaces older ones with[tool response truncated - call this tool again]on every request. Once the history holds more than N tool responses, each new one turns another retained response into the placeholder, which changes a message in the middle of the history. By the exact-prefix rule, the part of the history after that message cannot reuse the previous request's cache entry. That is an inference from how the request is built, not a measured miss. The docs present this as a context-size control and do not mention caching. - Context truncation. With the
Truncateoverflow action, the gateway removes older non-system messages until the conversation fits, which moves the start of the history. See pipelines. - Skill-based tool hiding. With
HideToolsWithoutSkill, the gateway filters the tool list it sends upstream. If that filtered list differs from one request to the next, so does the prefix; compare thetoolsfield of consecutive requests with the script above. The MCP tool-definition guide explains the flag. - Routing. A gateway that routes by complexity can choose a different model or reasoning effort per request; see Route by complexity, price by the task. For cache purposes, a different model is a different cache.
Two limits apply. The gateway forwards prompt_cache_key to the upstream provider for OpenAI-compatible requests, but I found no handling of Anthropic cache_control in the request code, so do not assume you can place explicit breakpoints through the gateway without testing it. And for non-managed providers, cache behavior is whatever that provider does.
If the numbers show a loss, the conservative changes are to raise ToolContextCount or leave it unset for short tool loops, avoid writing memory in the middle of a session you want to cache, and keep one model per conversation. Test each against your own usage logs, not against the sample numbers above. The learn guide on caching separates prompt caching from response caching, and the inference reference lists the request parameters.
How long does a cache entry last?
An entry that expired is a miss that no prefix hygiene fixes.
- OpenAI, GPT-5.6 and later:
prompt_cache_options.ttlsupports30m, the default. The entry stays eligible for 30 minutes after its latest write or reuse. - OpenAI, earlier models:
prompt_cache_retentionisin_memory(typically 5 to 10 minutes of inactivity, up to one hour) or24hon models that support it. The default depends on the organization's data-retention policy. - Anthropic: 5 minutes by default, refreshed on each use at no extra charge, or 1 hour at additional cost.
If the gap between turns in your product is longer than the window, because users read and think, a 1-hour or 24-hour setting can cost less than rewriting the cache on every turn. If turns arrive every few seconds, the default window already covers you and the cause is elsewhere.
FAQ
Why is my OpenAI prompt cache hit rate 0%?
Check the minimum cacheable length first (1,024 tokens on GPT-5.6 and later, and variable on earlier models, depending on tools, images, schemas, and reasoning effort). Then check the usual breakers above, then whether your requests arrive within the cache lifetime. OpenAI also notes that a shared prefix is not always a cached prefix: in implicit mode, a changing user message can be written to the cache while the stable part before it has no breakpoint of its own.
Does a lower hit rate after compaction mean something is wrong?
Not by itself. Compaction changes the prefix by design. Compare total input cost, not only the rate.
Do cached tokens count toward rate limits?
On OpenAI, yes: cached input tokens still count toward tokens-per-minute limits, so caching lowers cost and latency but not your rate-limit consumption.
Can I get the same numbers from the AIVAX API?
Yes, from usage.prompt_tokens_details.cached_tokens and usage.prompt_tokens on chat/completions, with the formula above. Whether a given upstream model reports cached tokens depends on the provider.