Prompt Caching Economics: Break-Even, TTL and Hit Rate
Cost & Pricing

Prompt Caching Economics: Break-Even, TTL and Hit Rate

Cache writes cost more than normal input, so caching only pays above a read threshold. Here is the arithmetic for TTL choice, hit rate and breakpoint placement.

Prompt caching is usually presented as a straightforward discount. It is not: writing to the cache costs more than an ordinary input token, and reading from it costs much less. Whether you come out ahead depends on how many reads follow each write.

That makes caching an arithmetic decision rather than a default. Here is the arithmetic, including the case where enabling caching makes your bill worse.

The three prices

Providers publish caching as multipliers on the base input rate. As of August 2026, Anthropic's published structure is:

uncached input        1.00x base
cache write, 5m TTL   1.25x base
cache write, 1h TTL   2.00x base
cache read            0.10x base

OpenAI publishes a comparable shape for GPT-5.6: cache reads retain roughly a 90% discount, with explicit cache writes billed at 1.25x normal uncached input. Sol's cached input rate is listed at $0.50 against $5 uncached.

Two structural facts follow. Reads are cheap enough that the discount is transformative when you get them. Writes carry a premium large enough that a cache you write and never read is a straight loss.

Break-even reads per write

Set the cached path equal to the uncached path for n total requests against the same prefix:

uncached: n x 1.00
cached:   1.25 + (n - 1) x 0.10      (5-minute TTL)

Solve for the crossover:

n x 1.00 = 1.25 + (n - 1) x 0.10
0.90n = 1.15
n = 1.28

So with a five-minute TTL, the second read already pays. Two requests against the same prefix within the window and you are ahead.

The one-hour TTL is a different calculation, because the write costs twice base:

n x 1.00 = 2.00 + (n - 1) x 0.10
0.90n = 1.90
n = 2.11

Three requests. Still a low bar, but not as low — and the choice between TTLs is really a choice about whether your traffic reliably delivers a second request inside five minutes.

Choosing the TTL from your inter-request gap

Measure the distribution of gaps between consecutive requests sharing a prefix. That single distribution answers the question.

If the p75 gap is under five minutes — an active agent loop, a live chat session, a busy API endpoint — the short TTL wins, because you get the cheap write and almost always hit.

If the p75 gap is between five minutes and an hour — bursty traffic, a user who reads a response then follows up, batch jobs spaced out across a shift — the long TTL is worth its doubled write cost, because the alternative is repeatedly paying 1.25x for writes that never get read.

Worked, for 100 requests a day against one prefix, with gaps averaging 20 minutes:

5m TTL:  almost every request misses and re-writes
         100 x 1.25 = 125 base-equivalents

1h TTL:  one write, 99 reads
         2.00 + 99 x 0.10 = 11.9 base-equivalents

no cache: 100 x 1.00 = 100 base-equivalents

The wrong TTL is 25% worse than not caching at all. The right one is eight times better. This is the specific failure mode worth guarding against.

Hit rate is what you actually control

The break-even numbers above assume the cache is hit. In practice the dominant variable is hit rate, and hit rate is a property of how you build prompts.

Effective input price as a function of hit rate h, on a short TTL:

effective = h x 0.10 + (1 - h) x 1.25
h = 0.95   ->  0.156x base
h = 0.80   ->  0.330x base
h = 0.50   ->  0.675x base
h = 0.20   ->  1.020x base
h = 0.10   ->  1.135x base

Below roughly 22% hit rate, caching costs more than not caching. That is the number to remember, and it is why "just turn caching on everywhere" is bad advice for endpoints whose prompts vary from the first token.

What destroys hit rate

Caching is a prefix match on exact bytes. Any change anywhere in the prefix invalidates everything after it. The usual culprits, in the order you will find them:

  • A timestamp or date in the system prompt. Changes every request, sits at position zero, invalidates the entire cache.
  • Session or user identifiers interpolated early. Produces a per-user cache with no cross-user sharing.
  • Non-deterministic serialisation. Serialising a dictionary or set without sorting keys produces different bytes for identical content.
  • A tool set that varies. Tools render before the system prompt, so adding, removing or reordering one invalidates everything.
  • Conditional system prompt sections. Every flag combination is a distinct prefix, fragmenting one cache into many.
  • Switching models mid-session. Caches are model-scoped; the switch is a full rebuild.

Do not diagnose these by reading code. Read the cache-read token count your provider returns in the usage object. If it is zero across repeated requests with what you believe is an identical prefix, diff the rendered bytes of two consecutive requests and the culprit will be obvious.

Where to place the breakpoint

The rule that follows from prefix matching: put the breakpoint at the end of the shared portion, not the end of the whole prompt.

A common mistake is to mark the end of a prompt whose tail varies per request. Every request then writes a distinct cache entry, nothing is ever read, and you pay the write premium on all of it — the worst possible outcome, and one that looks like caching is enabled.

Order content by stability: tool definitions and frozen system prompt first, session-stable material next, per-request content last. Then place the marker at the boundary between stable and volatile.

Two operational details worth knowing

Concurrent requests cannot read what is still being written. Fan out ten identical-prefix requests simultaneously and all ten pay full price. Send one, wait for it to begin streaming, then send the other nine.

There is a minimum cacheable prefix. Below it, nothing caches and no error is raised — you simply see zero cache-creation tokens. The threshold varies by model and is not monotonic across generations, so check for the specific model you are using rather than assuming.

Where caching does not help

It does nothing for output tokens, which for generation-heavy workloads is most of your bill. It does nothing for prompts that differ from the first token. And it does not remove the quadratic growth in agent conversations — it discounts the growing prefix, which is a large win, but the growth continues underneath.

If you are on flat-rate access, none of the cost arithmetic applies, and caching still matters: a cache read is faster than reprocessing the prefix, so the latency benefit persists whether or not you are paying per token.

Common questions

How many cache reads does it take before caching saves money?

With a write premium of 1.25x base and reads at 0.10x, the crossover is under two requests against the same prefix. A doubled write premium for a longer TTL pushes it to roughly three.

Can enabling prompt caching make my bill higher?

Yes. Below roughly a 22% hit rate the write premium exceeds the read savings. It happens most often when the cache breakpoint is placed after per-request content, so every request writes a fresh entry that is never read.

Should I use the short or the long cache TTL?

Measure the gap between consecutive requests sharing a prefix. If the p75 gap is under five minutes, take the cheaper write. If it falls between five minutes and an hour, the longer TTL is worth its doubled write cost.

Similar articles

Batching Requests to Save Money: When It Pays
Cost & Pricing
Cost & Pricing·8 min read

Batching Requests to Save Money: When It Pays

Batch endpoints trade latency for a large discount, and packing items into one prompt amortises overhead. Here is when each pays and when neither does.

Read
Cutting Token Usage Without Making Answers Worse
Cost & Pricing
Cost & Pricing·8 min read

Cutting Token Usage Without Making Answers Worse

Most prompts carry substantial waste. Nine techniques that reduce spend while usually improving output quality, ordered by how much they actually save.

Read
Forecasting AI Spend Without Guessing
Cost & Pricing
Cost & Pricing·8 min read

Forecasting AI Spend Without Guessing

Most AI budget forecasts are a headcount multiplied by a hopeful number. Here is a model that decomposes spend into drivers you can actually measure and control.

Read