Subagents and Delegation: When Splitting an Agent Helps
Subagents buy parallelism and a clean context window, and cost you shared understanding. Here is when that trade is worth making, and when it is not.
The pitch for subagents is intuitive. One agent is doing too much, its context is full of noise, and the work looks parallelisable. So you split it: an orchestrator that plans, and workers that each own a piece.
Sometimes that is exactly right. More often it converts a slow, legible failure into a fast, confusing one. The difference comes down to what kind of work you are splitting.
Subagents buy you exactly two things
Context isolation. A subagent that reads forty files and returns three paragraphs keeps thirty-seven files of noise out of the parent context. That is the real win, and it is bigger than most people expect. The parent stays coherent because it never sees the raw material.
Parallelism. Independent lines of enquiry can run at once, so wall-clock time drops even though total token spend rises. Anthropic described its research system as an orchestrator-worker design where a lead agent plans and spawns three to five subagents in parallel, then synthesises their findings.
That is the list. If your proposed split does not deliver one of those two, you have added coordination overhead for nothing.
What it costs
The cost is shared understanding. Anthropic reported that its multi-agent research system consumed roughly fifteen times the tokens of an ordinary chat interaction, with a single agent already at about four times. Fan-out is not free, and the multiplier is not marginal.
Worse than the tokens is the fragmentation. Cognition made this argument directly in Do Not Build Multi-Agents, with a worked example: split "build a Flappy Bird clone" into a background subtask and a bird subtask, and one worker misreads its brief and produces Super Mario Bros scenery. Neither worker is wrong given what it was told. The system is wrong because the brief was lossy.
Every delegation boundary is a compression step. You cannot hand a subagent everything the parent knows, so you hand it a summary, and the summary is where the misunderstanding lives.
The rule that survives contact with production
Split reads. Do not split writes.
Reading is naturally parallel and naturally lossy-tolerant. Three subagents searching three codebases cannot conflict with each other, and if one returns something useless the parent can ignore it. The worst case is wasted tokens.
Writing is neither. Two agents editing the same repository make decisions that must agree — naming, structure, interfaces, assumptions about what the other one did. They have no shared context in which to agree. Cognition later refined its position to roughly this: multi-agent setups work best when writes stay single-threaded and the extra agents contribute intelligence rather than actions.
That is a useful phrasing to keep. Subagents that gather, verify, critique and summarise are cheap to get right. Subagents that mutate shared state are where the wheels come off.
Designing the handoff
Treat a subagent brief as an API contract, not a chat message. Vague briefs are the single largest source of wasted subagent work.
- State the objective in output terms. "Return the file path and line number where the retry limit is configured" beats "look into the retry logic".
- Include the constraints the parent knows. Which directories are in scope, which approaches were already ruled out, what the parent has already tried.
- Specify the return shape. A fixed structure the parent can parse beats prose it has to re-read. This is where structured output earns its keep.
- Give a budget. Maximum steps or maximum tool calls, so a confused worker fails fast rather than grinding.
- Say what "not found" looks like. Without an explicit escape hatch, a subagent will invent a plausible answer rather than return nothing.
Orchestrator patterns worth using
Fan-out research, single writer. Workers gather evidence in parallel; the parent alone decides and edits. This is the safest shape and covers most real use cases.
Reviewer subagent. A second agent with a fresh context checks the first one against the original requirements. Fresh context is the point — it has not been persuaded by the reasoning that produced the work.
Specialist with a narrow toolset. A subagent that only has database tools cannot wander into the filesystem. Delegation doubles as a privilege boundary, which matters when the work touches untrusted content.
Cheap tier for mechanical steps. Extraction, formatting and classification do not need your strongest model. Routing those to a smaller one is one of the few cost levers that does not degrade the result.
Try these before you split
Most "we need subagents" problems are context management problems wearing a hat. Before adding an orchestrator, try:
- Better tools. A tool that returns ten relevant lines instead of a whole file removes the noise you were trying to isolate away.
- Summarising in the loop. Compact old turns in place rather than spawning a process to keep them out.
- Sequential phases. Research, then plan, then execute — with an explicit context reset between phases. You get isolation without concurrency.
- Just a bigger budget. Sometimes one agent with more steps beats three agents with a coordination problem.
A decision rule
Delegate when all four hold: the subtask is independently verifiable, it returns far less than it consumes, it does not write to state another agent touches, and its brief fits in a paragraph without hand-waving.
If any of those fails, keep it in one loop. A single agent that takes twice as long is easier to debug than a fleet that disagrees with itself, and debuggability is what you will actually spend your time on.
Common questions
Do subagents make an agent more reliable?
Only for read-heavy work. Isolating research into subagents keeps the parent context clean, which helps. Splitting work that writes to shared state usually hurts, because each boundary compresses away context the workers needed to agree with each other.
How much more do multi-agent systems cost?
Substantially. Anthropic reported its multi-agent research system used roughly fifteen times the tokens of a chat interaction, against about four times for a single agent. Fan-out multiplies both the prompt overhead and the number of loops.
How many subagents should an orchestrator spawn?
Few, and only for genuinely independent branches. Anthropic described spawning three to five in parallel for research tasks. Beyond that, synthesis becomes the bottleneck and the parent starts losing track of what each worker actually returned.