Cutting Token Usage Without Making Answers Worse
Most prompts carry substantial waste. Nine techniques that reduce spend while usually improving output quality, ordered by how much they actually save.
Token reduction has a reputation as a quality trade-off. Frequently it is the opposite: the same changes that cut spend also remove noise the model was struggling through. Here is what works, in rough order of impact.
1. Truncate tool output
In agent workloads this is the biggest single lever, by a distance. A search returning 1,200 matches, a file read returning 4,000 lines, a verbose test log — each lands in context and stays there for the rest of the session.
Cap results and say you did: "showing first 50 of 1,203 matches". The model handles truncation gracefully when it knows it happened.
2. Select files instead of dumping directories
Three relevant files beat thirty. This cuts cost and improves accuracy simultaneously, because irrelevant material actively degrades attention.
3. Use prompt caching
Where supported, a stable prefix — system prompt, tool definitions, fixed reference material — can be cached at a substantial discount. The requirement is that the prefix stays byte-identical, so keep timestamps, session IDs and anything else volatile out of it.
Order matters: static content first, dynamic content last. A single changing character at the top invalidates everything after it.
4. Compact conversation history
Long sessions carry their entire transcript. Replace old turns with a summary that preserves the objective, decisions made, and approaches already tried. Compact at 60–70% of the window rather than waiting for the limit.
5. Cap output length
Output tokens usually cost several times input tokens. Ask for what you need — "answer in under 100 words", "return only the diff" — and set max_tokens as a backstop.
A great deal of spend goes on preamble and summary nobody reads.
6. Route cheap work to cheap models
Commit messages, renames, docstrings and simple test scaffolding do not need a frontier model. Static routing by task type captures most of the available savings with almost no machinery.
7. Trim the system prompt
System prompts accumulate. Rules get added after each incident and are never removed. Every one is re-sent on every request, forever.
Read yours. Delete anything that is aspirational, duplicated, or no longer true. Long prompts also degrade adherence, so this often improves behaviour too.
8. Stop paying for abandoned generations
When a user navigates away, abort the upstream request. Without cancellation, generation runs to completion and you pay for output nobody sees.
9. Batch where latency allows
For offline work — classification, bulk summarisation — batch endpoints where offered are considerably cheaper in exchange for delayed results. If nobody is waiting, take the discount.
Watch for accidental resends
A recurring and expensive bug: something large gets rebuilt into the prompt on every request that did not need to change. A full schema dump, a directory listing, a config file serialised into the system prompt.
It is invisible in code review because each individual call looks reasonable. It shows up immediately in per-request token logs as a suspiciously constant floor. Anything static should be cached; anything unnecessary should be removed entirely.
Retries are silent spend
Every retry pays full input cost again. An agent looping three times on a failing approach costs four times the successful path, and nothing in your metrics flags it unless you look.
Track turns per completed task alongside tokens. A rising average usually means degraded tool descriptions or a model that has started struggling with something — both fixable, and both invisible if you only watch total spend.
What not to do
- Do not strip whitespace from code. Savings are trivial and comprehension suffers.
- Do not remove examples. One good example usually pays for itself several times over in fewer retries.
- Do not compress prompts into terse shorthand. Clarity is worth more than the tokens it costs.
- Do not skip the reasoning field in structured outputs. Removing it saves a little and costs accuracy.
Measure before optimising
Log input and output tokens per request, grouped by feature. The distribution is almost always surprising — a single endpoint nobody thought about is frequently the majority of spend. Optimise that one, ignore the rest.
And note that if you are on flat-rate access, most of this stops mattering for cost and continues to matter for quality and latency. That is a reasonable trade to make deliberately.
Common questions
What is the biggest source of wasted tokens?
In agent workloads, untruncated tool output. Search results, file reads and test logs enter context in bulk and are carried on every subsequent turn.
Does prompt caching require anything special?
The cached prefix must stay byte-identical, so put static content first and anything volatile such as timestamps last. A single changed character invalidates the rest.
Will shortening prompts make answers worse?
Usually the opposite. Removing irrelevant material improves attention. What does hurt is deleting examples or reasoning steps, which typically costs more in retries than it saves.