Context Window Cost Tradeoffs: Bigger Is Not Free
A million-token window changes what you can do and what you will pay. Here is the arithmetic for filling one, and when retrieval beats stuffing on cost.
Large context windows are commonly discussed as a capability: how much can you fit. The more consequential question for most teams is what it costs to keep filling one, turn after turn, when the model is stateless and every turn resends the lot.
Several current models list a one-million-token context window — Kimi K3, GLM 5.2 and DeepSeek V4 Pro among open models, and Anthropic lists 1M for its current Claude models. Being able to use it and being able to afford to use it are different questions.
The cost of one full window, once
Start with the simplest calculation. At $5 per million input tokens — Anthropic's listed Claude Opus 5 rate as of August 2026 — a single fully packed one-million-token request costs $5 in input alone.
That is a striking number on its own, but the important part is what happens next. Send a follow-up question against the same context and you pay $5 again, because nothing is retained between requests. Ten turns of conversation against a full window is $50 of input, for what feels to the user like one document and ten questions.
Compare against a 50,000-token context:
50k window, 10 turns: 10 x 0.05 x $5 = $2.50
1M window, 10 turns: 10 x 1.00 x $5 = $50.00
Twenty times the cost, for context that is often twenty times less relevant per token.
Then add the growth term
Real sessions do not hold context constant. They accumulate. If a session starts at some baseline and adds per turn, cumulative input over n turns is:
cumulative_input = n x baseline + growth x n x (n - 1) / 2
The second term is quadratic, which is why long sessions get expensive faster than people expect. A session starting at 100k and adding 20k per turn:
10 turns: 1,000k + 20k x 45 = 1,900k tokens -> $9.50
20 turns: 2,000k + 20k x 190 = 5,800k tokens -> $29.00
40 turns: 4,000k + 20k x 780 = 19,600k tokens -> $98.00
Doubling the session length quadruples the cost. This is the single most important shape to internalise about long-context work, and it is the reason compaction pays for itself so quickly.
Long-context pricing tiers
Some providers charge a premium above a context threshold rather than a flat rate across the window. OpenAI's published GPT-5.6 Sol pricing, for example, rises from $5 input and $30 output per million to $10 and $45 for long-context requests. Anthropic currently lists its 1M window at standard rates with no long-context premium.
Where a tier boundary exists, it creates a step function in your cost curve. A workload that averages just below the threshold and occasionally crosses it will show cost jumps that look inexplicable in an averaged dashboard. Check whether your provider has one, and where.
Caching changes the arithmetic, conditionally
If the large part of your context is stable across requests, caching transforms this. Anthropic prices cache reads at roughly 10% of the base input rate, with cache writes at 1.25x for the five-minute TTL and 2x for the one-hour TTL. OpenAI publishes a comparable structure for GPT-5.6, with Sol cached input at $0.50 against $5 uncached.
Applied to the 1M-window example, with the document cached and only the question varying:
first request (cache write): 1.0M x $5 x 1.25 = $6.25
each subsequent read: 1.0M x $5 x 0.10 = $0.50
10 turns cached: $6.25 + 9 x $0.50 = $10.75
10 turns uncached: $50.00
A four-and-a-half-fold reduction. The condition is strict: the cached prefix must stay byte-identical, and it must be at the front. A timestamp, a session ID, or a per-request preamble at the top of the prompt invalidates everything after it and you pay full price while believing you are caching. Verify against the cache-read token counts your provider reports rather than assuming.
Retrieval versus stuffing: where the line sits
The choice is not ideological. It is a break-even, and the break-even depends on how many turns you expect and whether the context is cacheable.
Retrieval costs you an embedding or search step and returns a much smaller context. Stuffing costs you nothing extra in engineering and returns a much larger one. Set them equal:
stuffing: turns x full_context x input_price
retrieval: turns x (retrieved_size x input_price) + turns x retrieval_cost
With a 400k-token corpus, retrieval narrowing it to 20k, negligible retrieval cost, and $5 per million:
stuffing, 8 turns: 8 x 0.4 x $5 = $16.00
retrieval, 8 turns: 8 x 0.02 x $5 = $0.80
Twenty-fold. Now cache the stuffed version, which is the fair comparison when the corpus is stable:
cached stuffing: (0.4 x $5 x 1.25) + 7 x (0.4 x $5 x 0.10) = $2.50 + $1.40 = $3.90
Still worse than retrieval, but the gap has narrowed from twenty-fold to five-fold — and stuffing requires no retrieval infrastructure, no chunking strategy, and no failure mode where the right passage is not retrieved. For many teams that is a reasonable trade at this margin.
The general rule that falls out: retrieval wins on cost when the corpus is much larger than what any one question needs, the sessions are long, and the corpus changes often enough to break caching. Stuffing wins when the corpus is stable, moderate in size, and the failure cost of missing a relevant passage is high.
Quality is not monotonic in context size
A cost article should still say this, because it changes the decision. Filling a large window with marginally relevant material does not simply cost more — it frequently produces worse answers, because irrelevant content competes for attention with the material that matters.
So the choice between a tight 30k context and a sprawling 500k one is rarely a straightforward quality-versus-cost trade. Often the cheaper option is also the better one, and the expensive option is expensive precisely because it was assembled without selection.
Practical rules
- Compact at 60–70% of the window, not at the limit. Waiting until you must compact means paying peak cost for many turns first.
- Put stable content first and volatile content last, so caching has something to hold onto.
- Measure cumulative session input, not per-request input. The quadratic term is invisible in a per-request average.
- Cap tool output before it enters context. In agent workloads this is the largest single driver of window growth.
- Check for a long-context pricing tier and, if one exists, whether your traffic straddles it.
If you are on flat-rate access, the cost half of this stops applying and the quality and latency halves do not. Large contexts remain slower and often less accurate, which is reason enough to keep them tight.
Common questions
Does a bigger context window cost more per token?
Not always per token, but some providers add a long-context tier above a threshold. The bigger effect is volume: a stateless model resends the whole context every turn, so a large window multiplies cost by turn count.
Is retrieval cheaper than putting everything in the context window?
Usually, but the gap narrows sharply once caching is applied to the stuffed version. Compute both against your own corpus size and expected turn count before building retrieval infrastructure you may not need.
When should I compact a long conversation?
At roughly 60-70% of the window rather than at the limit. Because cumulative input grows with the square of turn count, waiting until you are forced to compact means paying peak per-turn cost for many turns beforehand.