Prompt Caching: Stop Paying for the Same Tokens Twice
Fundamentals

Prompt Caching: Stop Paying for the Same Tokens Twice

Prompt caching reuses the prefill work for a repeated prefix, cutting cost and first-token latency. The whole technique comes down to one ordering rule.

Every request you send is processed from scratch. If your agent sends a 12,000-token system prompt and tool schema on turn one, then again on turn two, the provider does the same prefill work twice and bills you twice.

Prompt caching removes that duplication. It is one of the few optimisations that improves cost and latency at the same time without changing model behaviour at all, and most of the work is arranging your prompt in the right order.

What is actually cached

Not the answer. Prompt caching stores the intermediate state — the key/value tensors — produced while processing a prefix of your input. On a hit, the provider skips recomputing that prefix and starts work from where the cache ends.

Two consequences follow. The output is still generated fresh, so caching does not make responses repetitive or stale. And because it is a prefix cache, a change anywhere early in the prompt invalidates everything after it.

That second point is the entire design constraint. There is no partial credit for a prompt that differs in the first hundred tokens and matches for the next twenty thousand.

The one rule

Order your prompt from most stable to least stable.

  1. System prompt and persona
  2. Tool definitions
  3. Static domain context, style guides, schemas
  4. Retrieved documents for this session
  5. Conversation history
  6. The current user message

Anything that varies per request belongs at the end. The classic own-goal is injecting the current timestamp, a request ID or a randomly ordered retrieval set at the top of the system prompt — that alone can take your hit rate to zero while everything looks perfectly reasonable in code review.

Automatic versus explicit

The two major API styles differ in who decides what to cache.

OpenAI caches automatically. Prompts at or above a minimum length are eligible with no code change, cache hits require an exact repeated prefix, and matching proceeds in increments after the minimum. You do not mark anything; you just keep your prefix stable and watch the cached token count in the usage response.

Anthropic is explicit. You place cache_control markers on content blocks to declare cache breakpoints, with a small fixed number of breakpoints available per request. More work, more control — you decide exactly where the cached region ends.

Both have a minimum cacheable length that varies by model, and content below it is silently not cached. Check the current figure for the model you are calling rather than assuming; it has changed more than once.

The economics

Cached reads are dramatically cheaper than fresh input. OpenAI has offered a 50 percent discount on cached input tokens since the feature launched, with steeper discounts on newer model families. Anthropic prices cache reads at roughly 10 percent of the base input rate.

Writes are where the nuance is. Anthropic charges a premium to write a cache entry — about 1.25x base input for the five-minute TTL, and roughly 2x for the one-hour option. OpenAI historically charged no write premium, though newer model families such as GPT-5.6 and later price cache writes at 1.25x the uncached input rate.

So caching is not free money on the first call. With a 1.25x write and a 0.1x read, two uses of a cached prefix average out to about 0.675x — already a saving, but only from the second hit onward. A prefix used once per hour with a five-minute TTL costs you more than no caching at all.

What breaks a cache

  • Timestamps and IDs in the system prompt. The most common cause by a wide margin.
  • Reordered tool definitions. Serialising tools from an unordered map produces a different byte sequence per process.
  • Editing history. Summarising or trimming old turns rewrites the middle of the prompt and invalidates everything downstream.
  • Non-deterministic JSON serialisation. Key order, whitespace and float formatting all count.
  • Expiry. Anthropic caches default to a five-minute TTL that refreshes on each use; OpenAI entries are evicted after a period of inactivity. Traffic gaps mean cold starts.
  • Model or parameter changes. A different model, or in some cases different sampling configuration, is a different cache.

Agents are the ideal case

Agent loops benefit more than anything else, because the conversation is append-only by nature. Each iteration adds a tool result and a model turn to the end while the expensive front — system prompt, tool schemas, project context — stays byte-identical.

Get this right and a twenty-iteration session pays full price for its prefix once. Get it wrong and it pays twenty times, which for a large tool schema is the dominant line on the bill.

The tension is compaction. Rewriting history to control context growth is exactly the operation that destroys a prefix. The practical compromise is to compact rarely and in large steps rather than trimming a little each turn, so you take one cache miss instead of twenty.

Verify it, do not assume it

Every provider reports cached token counts in the usage block of the response. Log them. Compute hit rate as cached input tokens over total input tokens and put it on a dashboard next to cost and TTFT.

Then alert on it. A prompt refactor that quietly moves one dynamic field upward produces no errors, no test failures and no behaviour change — just a bill that doubles and a first-token latency that gets noticeably worse. Hit rate is the only signal that catches it.

When it does not help

Short prompts below the minimum length never cache. Genuinely unique one-shot requests — a batch job over a million distinct documents with no shared preamble — have no reusable prefix worth the write premium. And low-traffic endpoints with long gaps between calls will keep paying writes for entries that expire before a second use.

In those cases the honest answer is that caching is not your optimisation. Reducing input tokens or choosing a cheaper model is.

Common questions

Does prompt caching make responses less varied?

No. Only the prefill state for the prompt prefix is reused. Generation still runs normally, so sampling behaviour and output variety are unchanged.

Why is my cache hit rate zero?

Almost always something dynamic near the front of the prompt: a timestamp, a session ID, or tool definitions serialised in a different order each process. Diff two consecutive request bodies.

Is caching always cheaper?

No. Writes can cost more than a normal input token, so a prefix used only once before it expires costs more than not caching. It pays off from the second hit onward.

Similar articles

Batch Size and Throughput: The Trade Behind Every Token Price
Fundamentals
Fundamentals·9 min read

Batch Size and Throughput: The Trade Behind Every Token Price

Batching is why per-token prices are low and why latency varies under load. How it works, where it stops helping, and what it means for your requests.

Read
Grouped-Query Attention: Why Long Context Got Affordable
Fundamentals
Fundamentals·9 min read

Grouped-Query Attention: Why Long Context Got Affordable

GQA shrinks the KV cache by sharing keys and values across query heads. What it costs in quality, and why it is on the spec sheet of nearly every model.

Read
KV Cache Explained: The Memory Behind Long Context
Fundamentals
Fundamentals·8 min read

KV Cache Explained: The Memory Behind Long Context

The KV cache is why generation is fast and why long context is expensive in memory rather than compute. What it stores and what it costs you.

Read