What It Costs to Run a Fleet of Agents
Fan-out is the dominant cost term once you run agents in parallel. Here is the arithmetic for coordinator and subagent spend, plus what actually reduces it.
One agent working on one task has a cost you can reason about. A coordinator delegating to eight subagents, each of which re-establishes its own context, does not — at least not by intuition. The multiplier is larger than people expect and it compounds in a specific, predictable way.
This is the arithmetic for that multiplier, and the handful of design decisions that actually change it.
The fan-out term dominates everything else
A single-agent task costs roughly the accumulated context across its turns. A fan-out task costs that, plus the full independent context of every subagent, plus the coordinator re-reading every report that comes back.
Written out:
total = coordinator_context_growth
+ sum over subagents of (setup + own_context_growth + report)
+ coordinator cost of reading N reports
The third term is the one people forget. Each subagent report lands in the coordinator's context and stays there for the remainder of the run, so a coordinator that reads eight reports of 3k tokens each is carrying 24k extra tokens on every subsequent turn — not once.
A worked fan-out
Take a repository-wide audit: one coordinator, eight subagents, each assigned an independent module.
Coordinator setup (system prompt, tools, repo map): 12,000 tokens
Coordinator turns before delegating: 4 turns
Subagent setup, each (own prompt, tools, brief): 9,000 tokens
Subagent working turns, each: 6 turns, +2,000/turn
Subagent report back: 3,000 tokens
Coordinator turns after reports: 5 turns
Subagent input cost, per subagent, is the setup re-sent on each of six turns plus the accumulating work:
turn 1: 9,000
turn 2: 11,000
turn 3: 13,000
turn 4: 15,000
turn 5: 17,000
turn 6: 19,000
per subagent: 84,000 input tokens
x 8 subagents: 672,000
The coordinator carries its own 12k setup across nine turns, and from turn five onward it is also carrying the reports as they arrive. Cumulatively that is roughly 190,000 input tokens. Total input for the run: about 860,000 tokens.
At $5 per million input tokens — Anthropic's listed Claude Opus 5 rate as of August 2026 — that is around $4.30 in input alone, for one audit. Output adds perhaps another dollar. The comparable single-agent version of this task, working through modules sequentially with compaction, typically lands between a third and a half of that.
Fan-out is not free parallelism. It is a latency purchase, paid for in tokens.
The three cost terms you can actually move
Subagent setup, multiplied by N. Every subagent pays for its own system prompt and tool definitions, and pays again on every turn it takes. A 9k setup across six turns per subagent, eight subagents, is 432k tokens of pure repetition. This is the single largest addressable line, and it responds to two things: a leaner subagent prompt, and prompt caching where the setup is byte-identical across subagents.
Report size. Reports enter the coordinator's context permanently. A structured 500-token report instead of a 3,000-token narrative cuts the coordinator's carried weight by five-sixths. Specify the report format in the brief; do not leave it to the subagent's judgement.
Turn count per subagent. Because input grows with the square of turn count, halving turns per subagent cuts more than half the cost. Turn count usually falls when the brief is more complete up front — a subagent that has to discover its own scope spends turns doing it.
What does not move the number much
Trimming whitespace, shortening variable names in supplied code, or compressing prose into shorthand. These feel like optimisation and save a percent or two while making the model's job harder.
Similarly, routing subagents to a cheaper model helps only if the work genuinely fits — and a cheaper model that needs nine turns where the expensive one needed five can easily cost more, because turns are quadratic and price per token is linear. Measure this rather than assuming.
The costs that only appear at fleet scale
Failed and abandoned runs. At one agent, a failure is noticed. At fifty concurrent agents, a subagent that loops on a broken tool for its full turn budget is invisible unless you instrument for it. Track turns-per-completed-task as a first-class metric; a rising average is almost always a broken tool description or a degraded prompt.
Idle-but-billed orchestration. Nothing bills while an agent waits for a tool, but everything bills again when the coordinator re-reads state after the wait. Long tool calls in a fan-out therefore cost more than their latency suggests.
Retry storms. A transient upstream error that triggers retries across a whole fleet simultaneously multiplies one incident by your concurrency. Cap retries, back off exponentially, and put a spend ceiling on the key the fleet uses.
Deciding whether to fan out at all
The honest decision rule: fan out when the subtasks are genuinely independent and sizeable, and when wall-clock latency matters more than token spend. Do not fan out to parallelise something one agent could finish in a handful of tool calls — the setup cost per subagent swamps the work.
A useful threshold: if a subagent's expected work is smaller than its own setup cost, the delegation loses money. With a 9k setup, that means anything under roughly five substantive turns is better done inline.
Budgeting a fleet
Fleet spend is driven by concurrency and duty cycle, not headcount:
monthly = runs_per_day x avg_cost_per_run x 30
Measure avg_cost_per_run from real runs rather than from the arithmetic above, because the distribution has a long right tail and the mean sits well above the median. Then set the spend cap at roughly twice the resulting figure, on the fleet's own key, separate from human developer keys.
If your fleet is running continuously against a metered plan, flat-rate access is worth pricing — this is the usage pattern where it most reliably wins, since the fleet does not ration itself. If the fleet only runs nightly for twenty minutes, it does not, and metered is the correct answer.
Common questions
Why does adding subagents cost more than proportionally more?
Each subagent re-pays its own setup on every turn, and every report it returns is carried in the coordinator context for the rest of the run. You pay for N independent contexts plus a growing coordinator context, not one shared one.
When is fanning out to subagents worth the cost?
When subtasks are genuinely independent, each is substantial enough to exceed its own setup cost — roughly five or more turns — and wall-clock latency matters more than spend. Splitting a small job across agents reliably loses money.
How do I stop a runaway agent fleet from producing a huge bill?
Put a hard spend cap on the key the fleet uses, cap retries with exponential backoff, and alert on turns-per-completed-task rather than only on total spend. A rising turn average is the earliest signal something is looping.