Cutting Token Usage Without Making Answers Worse
Cost & Pricing

Cutting Token Usage Without Making Answers Worse

Most prompts carry substantial waste. Nine techniques that reduce spend while usually improving output quality, ordered by how much they actually save.

Token reduction has a reputation as a quality trade-off. Frequently it is the opposite: the same changes that cut spend also remove noise the model was struggling through. Here is what works, in rough order of impact.

1. Truncate tool output

In agent workloads this is the biggest single lever, by a distance. A search returning 1,200 matches, a file read returning 4,000 lines, a verbose test log — each lands in context and stays there for the rest of the session.

Cap results and say you did: "showing first 50 of 1,203 matches". The model handles truncation gracefully when it knows it happened.

2. Select files instead of dumping directories

Three relevant files beat thirty. This cuts cost and improves accuracy simultaneously, because irrelevant material actively degrades attention.

3. Use prompt caching

Where supported, a stable prefix — system prompt, tool definitions, fixed reference material — can be cached at a substantial discount. The requirement is that the prefix stays byte-identical, so keep timestamps, session IDs and anything else volatile out of it.

Order matters: static content first, dynamic content last. A single changing character at the top invalidates everything after it.

4. Compact conversation history

Long sessions carry their entire transcript. Replace old turns with a summary that preserves the objective, decisions made, and approaches already tried. Compact at 60–70% of the window rather than waiting for the limit.

5. Cap output length

Output tokens usually cost several times input tokens. Ask for what you need — "answer in under 100 words", "return only the diff" — and set max_tokens as a backstop.

A great deal of spend goes on preamble and summary nobody reads.

6. Route cheap work to cheap models

Commit messages, renames, docstrings and simple test scaffolding do not need a frontier model. Static routing by task type captures most of the available savings with almost no machinery.

7. Trim the system prompt

System prompts accumulate. Rules get added after each incident and are never removed. Every one is re-sent on every request, forever.

Read yours. Delete anything that is aspirational, duplicated, or no longer true. Long prompts also degrade adherence, so this often improves behaviour too.

8. Stop paying for abandoned generations

When a user navigates away, abort the upstream request. Without cancellation, generation runs to completion and you pay for output nobody sees.

9. Batch where latency allows

For offline work — classification, bulk summarisation — batch endpoints where offered are considerably cheaper in exchange for delayed results. If nobody is waiting, take the discount.

Watch for accidental resends

A recurring and expensive bug: something large gets rebuilt into the prompt on every request that did not need to change. A full schema dump, a directory listing, a config file serialised into the system prompt.

It is invisible in code review because each individual call looks reasonable. It shows up immediately in per-request token logs as a suspiciously constant floor. Anything static should be cached; anything unnecessary should be removed entirely.

Retries are silent spend

Every retry pays full input cost again. An agent looping three times on a failing approach costs four times the successful path, and nothing in your metrics flags it unless you look.

Track turns per completed task alongside tokens. A rising average usually means degraded tool descriptions or a model that has started struggling with something — both fixable, and both invisible if you only watch total spend.

What not to do

  • Do not strip whitespace from code. Savings are trivial and comprehension suffers.
  • Do not remove examples. One good example usually pays for itself several times over in fewer retries.
  • Do not compress prompts into terse shorthand. Clarity is worth more than the tokens it costs.
  • Do not skip the reasoning field in structured outputs. Removing it saves a little and costs accuracy.

Measure before optimising

Log input and output tokens per request, grouped by feature. The distribution is almost always surprising — a single endpoint nobody thought about is frequently the majority of spend. Optimise that one, ignore the rest.

And note that if you are on flat-rate access, most of this stops mattering for cost and continues to matter for quality and latency. That is a reasonable trade to make deliberately.

Common questions

What is the biggest source of wasted tokens?

In agent workloads, untruncated tool output. Search results, file reads and test logs enter context in bulk and are carried on every subsequent turn.

Does prompt caching require anything special?

The cached prefix must stay byte-identical, so put static content first and anything volatile such as timestamps last. A single changed character invalidates the rest.

Will shortening prompts make answers worse?

Usually the opposite. Removing irrelevant material improves attention. What does hurt is deleting examples or reasoning steps, which typically costs more in retries than it saves.

Similar articles

Prompt Caching Economics: Break-Even, TTL and Hit Rate
Cost & Pricing
Cost & Pricing·8 min read

Prompt Caching Economics: Break-Even, TTL and Hit Rate

Cache writes cost more than normal input, so caching only pays above a read threshold. Here is the arithmetic for TTL choice, hit rate and breakpoint placement.

Read
Cost Per Pull Request: A Unit Metric Worth Tracking
Cost & Pricing
Cost & Pricing·8 min read

Cost Per Pull Request: A Unit Metric Worth Tracking

Total AI spend tells you nothing actionable. Cost per merged pull request ties inference to delivered work and exposes exactly where the money goes.

Read
Forecasting AI Spend Without Guessing
Cost & Pricing
Cost & Pricing·8 min read

Forecasting AI Spend Without Guessing

Most AI budget forecasts are a headcount multiplied by a hopeful number. Here is a model that decomposes spend into drivers you can actually measure and control.

Read