Context Windows: Why More Is Not Always Better
Fundamentals

Context Windows: Why More Is Not Always Better

A million-token context window sounds like the end of chunking. In practice models get slower, pricier and less accurate long before you fill it. Here is why.

The context window is the total number of tokens a model can consider at once — your system prompt, the conversation so far, any files you pasted, and the response it is about to generate. All of it competes for the same budget.

Vendors advertise this number heavily because it is easy to compare. It is also one of the least useful numbers for predicting whether a model will do your job well.

Tokens are not words

A token is roughly a common chunk of characters. English prose averages about 0.75 words per token. Code is worse — punctuation, indentation and identifiers fragment badly, so a 500-line file often costs far more than its word count suggests.

Useful rules of thumb:

  • ~4 characters per token for English prose
  • ~3 characters per token for code
  • A 200-line source file: roughly 2,000–3,000 tokens
  • A dense technical PDF page: roughly 600–800 tokens

Non-English text is frequently far more expensive per character, because tokenizers are trained predominantly on English.

The window is shared with the output

This trips people up constantly. If a model has a 128k window and you fill 127k with context, there is no room left to answer. Most APIs will either truncate or error.

Always leave headroom for the response — and remember that in an agent loop the accumulated tool output is also consuming that budget on every turn.

Lost in the middle

The finding that matters most in practice: models do not attend uniformly across their context. Information at the very beginning and the very end is recalled reliably. Information buried in the middle of a long context is recalled markedly less well.

This is a robust, repeatedly-observed effect, and it has a direct consequence — where you put something matters as much as whether you included it.

Practically:

  • Put instructions at the start, or repeat them at the end.
  • Put the most relevant file closest to your question.
  • Do not assume that because something is "in context" the model has actually used it.

The "needle in a haystack" tests vendors publish measure retrieval of a single distinctive fact. That is a much easier task than reasoning over several facts scattered through a long document, which is what real work usually needs.

Cost and latency scale badly

Attention cost grows quadratically with sequence length. Implementations mitigate this, but the direction holds: doubling your context more than doubles the work.

You feel this as:

  • Time to first token climbing — the model must process everything before emitting anything.
  • Cost per call rising linearly with input tokens, on every turn of an agent loop.

Dumping your whole repository into context is technically possible and almost always the wrong move. You pay for every token, on every request, forever — and the model attends to the middle of it poorly anyway.

What to do instead

  1. Select, do not dump. Three relevant files beat thirty. This improves accuracy and cost simultaneously, which is rare.
  2. Put the question last. Recency is strongly weighted; end with what you actually want.
  3. Summarise long histories. In long agent sessions, compact old turns into a summary rather than carrying raw transcript.
  4. Use prompt caching where offered. A stable prefix — system prompt, tool definitions — can often be cached at a large discount. This requires the prefix to stay byte-identical, so keep volatile content like timestamps out of it.
  5. Measure, do not assume. Test whether adding context actually improved the answer. Frequently it does not.

Counting tokens before you send them

Guessing is how budgets get blown. Every major SDK exposes a tokenizer, and for rough planning you can count characters and divide. What matters is doing it before a request rather than discovering the cost afterwards.

Two habits pay for themselves quickly. First, log input token counts per request in development — you will find prompts far larger than you assumed, usually because something is being serialised into the system prompt on every call. Second, set a hard ceiling in your own code and fail loudly when a prompt exceeds it, rather than letting a runaway context silently multiply your bill.

Where the budget actually goes

In a typical agent session the breakdown is rarely what people expect. The system prompt and tool definitions are fixed and usually modest. The conversation history grows steadily. But the dominant cost is almost always tool output — file contents, search results, test logs — which arrives in large chunks and never leaves.

This is why truncating tool results is one of the highest-leverage optimisations available. Returning the first 50 matches instead of all 1,200 changes the trajectory of the entire session, because everything after that point carries the smaller payload forward.

When long context genuinely earns its keep

There are real cases: a single large document you must reason over holistically, a long debugging session where earlier state matters, or codebases where the relationships between files are the point. In those, a large window is the difference between possible and impossible.

The mistake is treating it as the default. Long context is a capability to reach for deliberately, not a substitute for deciding what matters.

Common questions

How many tokens is my codebase?

Divide total characters by about three for code. A 50,000-line project is very roughly 500k–750k tokens — which is why selective retrieval beats loading everything, even with a large window.

Does a bigger context window make a model smarter?

No. It changes how much the model can see, not how well it reasons. A smaller model with a huge window is still a smaller model, and accuracy typically degrades as the window fills.

Is RAG obsolete now that windows are large?

No. Retrieval is still cheaper, faster, and often more accurate because it puts less irrelevant material in front of the model. Long context and retrieval are complementary, not competing.

Similar articles

Context Length vs Effective Context: The Number You Can Use
Fundamentals
Fundamentals·9 min read

Context Length vs Effective Context: The Number You Can Use

A model advertising a one-million-token window does not reliably use one million tokens. The gap between the spec sheet and what actually works.

Read
Flash Attention Explained: Same Maths, Far Less Memory
Fundamentals
Fundamentals·9 min read

Flash Attention Explained: Same Maths, Far Less Memory

Flash attention makes long context practical by never writing the score matrix to memory. What it changes, what it does not, and where you feel the difference.

Read
How LLMs Actually Work: Next-Token Prediction, End to End
Fundamentals
Fundamentals·9 min read

How LLMs Actually Work: Next-Token Prediction, End to End

A large language model does one thing repeatedly. Here is what happens between your prompt and the first token out, and what that mechanism explains about behaviour.

Read