## Prompt caching

*How a model writes and what it costs*

*Last edited: 28 September 2026*

Prompt caching is the KV cache kept by the provider between requests. A request that starts exactly like a recent one skips prefill for that part: it pays a fraction of the input price and gets its first token sooner.

**In plain words:** A waiter who knows the regulars doesn’t ask again about your allergies and favourite table. But give a different name and they start from scratch. And they remember you for only a few minutes after your last visit.

*Interactive widget on the page: Pick a prompt layout. See how much of the second request hits the cache and what it costs.*

### How it works

- After a request, the provider keeps the prompt’s KV cache in blocks and indexes them by a hash of the prefix, that is, everything from the start of the prompt to the end of the block (see “The generation loop and KV cache”). A request that starts the same way loads those blocks instead of recomputing them.
- Only an identical prefix hits. Each token’s K and V depend on all tokens before it and, because of the causal mask, on nothing after it. So an appended suffix leaves the cache valid, while the first changed token invalidates everything after it, even if the rest is identical.
- The gain is twofold: cheaper input and a shorter time to first token, because prefill skips the cached part. It does not speed up generating the answer. The model sees exactly the same text, so the cache doesn’t change quality.

### Provider terms

- Anthropic (September 2026): reads at 0.1× the input price (even less on the newest models), writes at 1.25× for a 5-minute entry and 2× for a one-hour entry. You mark the cut points yourself (up to 4 `cache_control` markers) or set a single field for automatic mode. The minimum prefix is 512 to 4096 tokens depending on the model. A shorter one is silently not cached.
- OpenAI caches prefixes of 1024 tokens or more automatically. From GPT-5.6, writes cost 1.25×, reads 0.1×, and an entry lives for at least 30 minutes. Older models charge nothing extra for writes and by default keep an entry for about 30 minutes, up to 24 hours; with Zero Data Retention the default is 5–10 minutes of inactivity, up to an hour. Gemini from 2.5 has an automatic cache with no hit guarantee, and an explicit one billed per hour of storage.
- Every hit renews the entry’s lifetime. At Anthropic it counts from the start of the request, so a 4-minute generation uses up most of a 5-minute entry. With such an entry, a user who replies after a quarter of an hour starts with a write.

### How to lay out the prompt

- Stable parts first: tool definitions, system prompt, large documents, examples. Then the history, and the new message last. At Anthropic the hierarchy is `tools` → `system` → `messages`, so changing the tools invalidates everything.
- Only ever append to the history. An agent resends all of it every turn (see “The context window and agents”), and only an append-only history lets each turn read everything but its newest part from the cache. Editing an old message, compaction, cutting an old tool result or reordering tools invalidates the cache from that point. Saving window space (see “Context engineering and memory”) therefore has to be weighed against the cost of a miss.
- Typical prefix killers: a date and time or a session ID at the start, non-deterministic ordering of JSON keys or tools, switching the model, thinking level or response schema mid-conversation.
- Measure hits in `usage`: `cache_read_input_tokens` at Anthropic, `cached_tokens` at OpenAI. Hit rate is cached tokens divided by all input tokens. At Anthropic `input_tokens` counts only what follows the last breakpoint, so all input is `input_tokens` + `cache_creation_input_tokens` + `cache_read_input_tokens`; at OpenAI `input_tokens` already includes `cached_tokens`. Zero with a repeated prefix means something is changing it.
- Break-even: with writes at 1.25 and reads at 0.1, an entry pays for itself after a single hit (1.35 instead of 2 input prices). A one-hour entry at 2 needs two hits. A prefix that never repeats only pays the write surcharge.

### Check yourself

**Question:** How does prompt caching work, and how do you structure a prompt for it?

**Short answer:** It is the KV cache kept by the provider across requests. It only matches an identical prefix, because a token’s keys and values depend on everything before it: the first changed token invalidates the rest. So stable parts such as tools, the system prompt and large documents go first, history is append-only, and volatile data goes last. Reads usually cost 10% of the input price, writes can cost more than plain input, and entries live for minutes. The payoff is cheaper input and faster time to first token. The hit rate comes from the usage fields.

### Follow-up questions

- **The agent bill is high and the hit rate is low. What do you check?** Diff consecutive requests token by token and look for the first difference: a date or ID at the start, non-deterministic serialisation of tools and JSON, edited history, a change of model or thinking level. Then check whether the gaps between turns exceed the entry lifetime and whether the prefix meets the minimum length.
- **How is prompt caching different from a semantic cache?** Prompt caching stores the computation for an identical prefix, and the model still generates a new answer, so quality doesn’t change. A semantic cache returns an old answer to a similar question: it saves the whole call but may return the answer to a different question.
- **When does caching not pay off?** When the prefix rarely repeats within the entry’s lifetime. With a write surcharge, every entry that never gets a hit costs more than a request without caching.
- **Can the cache leak another user’s data?** Large providers don’t share the cache across organisations, but within your application users with the same prefix hit the same entry. A hit shows up as a shorter response time, so someone can test whether another user recently sent a given text. A 2025 audit found cache shared globally, across organisations, at 7 of 17 API providers, and at least five changed that only after disclosure. Isolate the cache per customer (a separate workspace at Anthropic; at OpenAI, from GPT-5.6, a separate prompt_cache_key, while on older models the key only steers routing) or keep sensitive data out of the shared prefix.

### Sources

- [Claude docs: prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
- [OpenAI docs: prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching)
- [Manus: Context Engineering for AI Agents](https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus)
- [Auditing Prompt Caching in Language Model APIs (arXiv)](https://arxiv.org/abs/2502.07776)

Interactive page: https://howaiworks.dev/prompt-caching/
