Why Agent Costs Are Unpredictable, and What to Do About It
Agent spend follows a heavy-tailed distribution, so the average is a poor planning number. Here is how to budget from percentiles and bound the tail instead.
Ask what a task costs and you will be given an average. Averages work fine when the underlying distribution is well behaved. Agent task costs are not well behaved: they are heavy-tailed, which means the mean sits well above the median and a small number of runs account for a large share of total spend.
That is not a flaw to be engineered away. It is a structural property of what agents do, and budgeting works much better once you accept it and plan around the shape rather than the centre.
Where the variance comes from
Four independent sources, each multiplicative rather than additive.
Turn count is not fixed and input grows with its square. Because the model is stateless, turn n resends everything from turns 1 to n-1. A task that takes 20 turns instead of 10 does not cost twice as much; it costs roughly four times as much. Turn count is decided by the model at runtime based on what it finds, so this multiplier is set after you have committed.
Tool output size is unbounded until you bound it. A search that returns 12 matches and one that returns 1,200 differ by two orders of magnitude in tokens, and the difference persists in context for every subsequent turn.
Retries and dead ends. An agent that tries an approach, discovers it fails, and backs out has paid full price for the excursion. Nothing distinguishes it from productive work in a token count.
Task difficulty is only known afterwards. Two tickets that read identically can differ tenfold in how much code needs to be read. The estimate you would have made before starting has no access to that information.
What the distribution actually looks like
Instrument per-task cost for a few hundred agent runs and the shape is consistent across teams. Illustrative figures from a coding-agent workload, in cost per completed task:
p10 $0.06
p50 $0.21
mean $0.58
p90 $1.40
p99 $6.80
max $31.00
Three things to read off that. The mean is nearly three times the median, which is the signature of a heavy tail. The p99 is thirty times the median. And a single run cost more than 150 typical runs put together.
Your own numbers will differ, but the ratios rarely do. If your measured mean-to-median ratio is close to 1, you are probably not measuring agent tasks — you are measuring something with a bounded turn count.
Why budgeting from the mean fails
Take that distribution and a plan to run 5,000 tasks a month:
budget from mean: 5,000 x $0.58 = $2,900
That is arithmetically correct as an expectation, and it will still be wrong most months, because the total is dominated by how many tail events happen to land in the period. With a tail this heavy, the month-to-month variance of the total is driven by a handful of runs.
The more useful framing splits the budget:
body: 5,000 x $0.21 (median) = $1,050
tail: the difference, which is variable = $1,850 expected
total expected = $2,900
plan for = $4,500 with headroom
Now you have two numbers that behave differently. The body is stable and predictable and you can forecast it. The tail is not, and the correct response to it is a bound rather than a forecast.
Bounding the tail is more effective than reducing the mean
This is the central practical point. Effort spent shaving 10% off typical runs moves the body. Effort spent capping the worst 1% moves the total more, because that 1% is a disproportionate share of it.
Cap turns per task. A hard ceiling — twenty-five turns, say — converts an unbounded loop into a bounded failure you can retry deliberately. Pick the ceiling from your own distribution: whatever value contains 98% of successful runs.
Cap context size. Stop the run, compact, or fail when accumulated context crosses a threshold. Runs that grow past your p95 context size rarely recover; they are usually lost and re-reading the same material.
Truncate every tool output. Not just the ones that look large. A single uncapped search result is the most common cause of a run leaving the body of the distribution.
Bound the task budget where the API supports it. Some providers accept a token budget the model is aware of, so it paces itself and finishes gracefully rather than being cut off mid-thought. That produces a better failure than a hard truncation.
Cap spend per key. The backstop that catches everything the other four missed.
Instrument for the tail specifically
Aggregate spend tells you a tail event happened, three weeks late. These tell you sooner:
- Turns per completed task, as a distribution rather than an average. A rising p90 with a flat median means something is degrading for a subset of tasks — usually a tool description or a changed prompt.
- Cost per task at p50, p90, p99, tracked over time. Watch the gap between them, not the levels.
- Abandonment rate. Tasks that hit the turn cap without completing are pure loss and should be a named metric.
- Cost attribution per task type. Heavy tails are rarely uniform. Usually one category of work — a particular repository, a particular kind of ticket — generates most of them, and that is actionable in a way that a global average is not.
Communicating this to people who want one number
Finance teams reasonably prefer a single figure. Give them one, framed correctly: "typical task $0.21, budget assumes $0.58 average, capped at $4,500 a month by a hard spend limit." That is one number for planning, one for expectation, and one for the worst case, and it holds up under questioning in a way that "about sixty cents a task" does not.
This is also the honest case for flat-rate access: it does not reduce the variance, it moves it onto someone else. For a workload with a heavy tail and continuous usage, that transfer is worth real money and it removes the forecasting problem entirely. For a workload that runs twice a week, it is not worth it — the tail is small in absolute terms and metered billing is cheaper. Which of those you are is answered by the same distribution you just measured.
The one thing not to do
Do not respond to variance by making the agent more timid — smaller context, fewer tools, tighter instructions to stop early. That reduces the tail by reducing capability, and the tasks you lose are disproportionately the difficult ones you most wanted help with. Bound the tail with hard limits and let the agent work freely inside them.
Common questions
Why is my average agent cost so much higher than a typical run?
Because the distribution is heavy-tailed. A small number of long runs, driven by high turn counts and large tool outputs, pull the mean well above the median. A mean-to-median ratio near three is normal for coding agents.
Should I budget from the average or the median?
Both, separately. Forecast the body of the distribution from the median times task count, treat the tail as a variance term you bound rather than predict, and set a hard spend cap at roughly 1.5x the total expectation.
What single control most reduces total agent spend?
A hard turn cap, set at whatever value contains about 98% of your successful runs. Because input grows with the square of turn count, capping the longest runs removes a disproportionate share of total spend.