Agent Cost Control: Patterns That Bound the Worst Case
AI Agents

Agent Cost Control: Patterns That Bound the Worst Case

Agent spend is not a per-call problem, it is a per-loop problem. Budgets, caching, tiering and circuit breakers that cap what a bad run can cost you.

Agent cost is not an average problem. The average run is fine. The problem is the run that took ninety-four steps because a tool kept returning a malformed response, and nothing in your system was empowered to stop it.

Which means most cost work should go into bounding the tail, not shaving the median. Prompt golfing saves a few per cent. A step cap saves the outlier that costs more than a thousand ordinary runs.

Understand where the money actually goes

Two structural facts drive agent economics.

First, models are stateless, so every step resends the whole accumulated transcript. Input tokens grow roughly quadratically across a run. Step twenty carries steps one through nineteen with it, which is why doubling the step count considerably more than doubles the cost.

Second, the dominant term is input, not output. Agents read enormously and write little. A run that produces a twenty-line patch might have ingested half a codebase to get there. Optimising output length is mostly wasted effort.

Together these say: the levers that matter are step count, context size per step, and the price you pay per input token.

Prompt caching is the highest-leverage lever

An agent loop is close to the ideal caching workload, because every step shares a long identical prefix with the previous one — system prompt, tool definitions, and the entire transcript so far.

The economics are favourable enough to change your architecture. On the Claude API, cache writes cost 1.25x base input for a five-minute TTL or 2x for one hour, and cache reads cost 0.1x base input. You pay a quarter extra once and then a tenth thereafter. Minimum cacheable prefix length varies by model, and you get up to four explicit cache breakpoints per request.

Three implications worth internalising:

  • Prefix stability is the whole game. Anything that changes early in the prompt invalidates everything after it. Put a timestamp in your system prompt and you have disabled caching for the entire run.
  • Order by volatility. Static system prompt first, then tool definitions, then stable context, then the moving transcript. Breakpoints go at the boundaries.
  • Do not rewrite history. Compaction that rewrites earlier turns invalidates the cached prefix, so a compaction pass has a real cost. Compact at deliberate boundaries, not continuously.

Verify it is working rather than assuming. Cached read and cached write token counts come back in the response; if reads are near zero, something in your prefix is unstable.

Budgets that are enforced, not requested

Telling a model to be economical does not bound anything. Budgets belong in the harness, checked before each call, in layers:

  1. Per-request ceiling. Maximum context and maximum output tokens for a single call. Catches a runaway tool result before it lands.
  2. Per-run budget. A token or currency ceiling for the whole task, decremented as you go. On exhaustion, stop cleanly and hand back partial state.
  3. Step cap. The simplest and most effective control there is. Pick a number from your own trace percentiles rather than intuition.
  4. Per-user and per-tenant windows. Prevents one caller consuming a shared capacity pool.
  5. Global circuit breaker. A spend rate that trips and degrades to a cheaper mode. Every serious incident of this kind was discovered on an invoice, not a dashboard.

Stall detection pays for itself

Most expensive runs are not doing expensive work. They are stuck, and the loop has no way to know it.

Two cheap detectors catch the majority:

Repeated identical calls. Same tool, same arguments, twice in one run. Hash the call and compare. On a repeat, inject an explicit observation saying the call was already made and giving the previous result — that alone frequently unsticks the model.

No state change over N steps. If nothing was written, no test outcome changed and no new file was read across five steps, the agent is circling. Stop and escalate.

Both are a few lines of code and both remove entire classes of runaway cost.

Tier your models by step, not by product

The instinct is to pick one model for the agent. The better shape is to pick per step, because agent runs contain a great deal of mechanical work — extraction, formatting, classification, summarising a tool result — that does not need frontier reasoning.

Two patterns are reliable:

Cheap by default, escalate on failure. Run the smaller model, validate the output, and retry on the stronger one only when validation fails. Treat expensive inference as the fallback rather than the default.

Small model as a filter. Use a cheap call to decide whether the expensive call is needed at all. Summarising a 40KB tool result down to the relevant lines before it enters the main context saves those tokens on every subsequent step, which is where the compounding works in your favour.

Measure this properly. Cost per resolved task is the metric, not cost per call. A cheaper model that needs three times the steps is more expensive, and per-call pricing hides that completely.

Design tools for token efficiency

Tool output goes into context and then gets resent on every subsequent step. A verbose tool is a recurring charge, not a one-off.

  • Return the relevant slice, not the whole file. A read tool with line ranges beats one without.
  • Paginate, with an explicit continuation token, rather than truncating silently.
  • Strip boilerplate — build logs, progress bars, repeated stack frames — before the result reaches the model.
  • Return errors that say what to do next. A useful error message costs a few tokens and saves several exploratory steps.
  • Prune superseded results from the transcript when a file has been re-read or a command re-run.

The billing model changes which controls you need

Worth being explicit, because it is often confused. Under metered per-token pricing, every control above is doing two jobs: preventing stalls and protecting the invoice. That dual purpose leads teams to set caps tighter than task quality wants, which shows up as agents that give up early.

Flat-rate access — what we sell, so weigh this accordingly — removes the second job. It does not remove the first. Step caps, stall detection and circuit breakers still matter, because a looping agent wastes wall-clock time and produces worse results regardless of who is metering the tokens. Predictable billing is a budgeting benefit, not a substitute for a bounded loop.

Where to start

  1. Instrument first: tokens, steps and cost per run, tagged by task type.
  2. Look at the ninety-fifth percentile run, not the median.
  3. Add a step cap at roughly twice that percentile.
  4. Turn on prompt caching and verify cache reads are actually happening.
  5. Add repeated-call and no-progress detection.
  6. Only then tune prompts and tool output sizes.

The first five are structural and hold as your prompts change. The sixth is maintenance work you will redo every time the system evolves, which is why it belongs last.

Common questions

Why do agents cost so much more than chat?

Because the model is stateless and each step resends the whole transcript, so input tokens grow roughly quadratically across a run. A single agent task can involve a dozen or more round trips, each carrying everything that came before it.

How much does prompt caching save on an agent loop?

A lot, because every step shares a long identical prefix. On the Claude API a cache write costs 1.25x base input for a five-minute TTL and a cache read costs 0.1x, so a stable prefix is read back at a tenth of the price on every subsequent step.

What is the single most effective agent cost control?

A step cap, derived from your own trace percentiles. Cost blowouts come from the tail of stuck runs rather than the median, and a hard ceiling on steps bounds the worst case in a way that prompt tuning never can.

Similar articles

Agent Failure Modes: A Taxonomy Worth Memorising
AI Agents
AI Agents·9 min read

Agent Failure Modes: A Taxonomy Worth Memorising

Agents fail in about eight recognisable ways, and each one needs a different fix. A field guide to spotting them from a trace and knowing what to change.

Read
Agent Memory: Managing Context Before It Manages You
AI Agents
AI Agents·8 min read

Agent Memory: Managing Context Before It Manages You

Long agent sessions fail because context fills with noise. Here are the practical strategies for deciding what an agent should remember and what to drop.

Read
Agent Observability: Tracing a Loop You Cannot Reproduce
AI Agents
AI Agents·9 min read

Agent Observability: Tracing a Loop You Cannot Reproduce

Agent failures are rarely reproducible, so logs are not enough. What to record per step, how to span a tool loop, and which metrics predict a bad run.

Read