Multi-Agent Systems: When They Help and When They Hurt
Two credible teams published opposite advice on multi-agent architectures within a day of each other. Both were right, because the answer depends on the task.
In June 2025 Cognition published "Don't Build Multi-Agents", arguing that splitting work across agents produces fragile systems built on conflicting assumptions. The following day Anthropic published how their multi-agent research system beat a single agent by 90.2 percent on internal research evaluations.
Neither was wrong. They were describing different tasks, and the difference between those tasks is the only thing you need to decide your own architecture.
The variable that decides it
Ask whether the work decomposes into parts that do not need to agree with each other.
Research does. Five subagents investigating five aspects of a question can work in isolation and their findings merge cleanly, because nothing one discovers invalidates another's approach. Anthropic's system is an orchestrator-worker design for exactly this shape: a lead agent plans, spawns parallel subagents, and synthesises what comes back.
Implementation does not. Cognition's example is a Flappy Bird clone split into "build the background with pipes" and "build the bird you can move". Subagent one drifted into a Super Mario Bros style background, subagent two built to a different visual assumption, and the pieces did not fit — not because either agent failed its brief, but because the briefs never encoded the shared assumptions that mattered.
Read-heavy exploration parallelises. Write-heavy construction on shared state does not.
The cost is not marginal
Anthropic reported that multi-agent systems use roughly 15 times more tokens than chat interactions, and that token usage alone explained about 80 percent of the performance variance in their research evaluations, with tool call count and model choice covering most of the rest.
That second finding deserves more attention than it usually gets. A large part of why multi-agent systems win is simply that they spend more compute on the problem. Before adding an orchestrator, check whether a single agent with a larger step budget, better tools and more thinking gets you the same place for less.
The economic conclusion follows: multi-agent architectures pay off when the task is valuable enough to justify a large token bill. For a high-stakes research report, easily. For a routine support reply, not remotely.
Patterns that work
- Orchestrator and workers for parallel search. Independent read-only investigations, findings returned as structured summaries, one agent synthesising. The canonical win.
- Context isolation for noisy exploration. Send a subagent to grep a large codebase and return ten relevant lines. The main agent gets the answer without the 50,000 tokens of noise. This is a context-management technique that happens to look like a multi-agent one.
- Sequential specialists. Draft, then review. The reviewer works from the artefact, not from the drafter's reasoning, and a fresh context often catches what the author cannot see.
- Separate models for separate cost profiles. A cheap model for mechanical extraction, a strong one for judgement. Routing by difficulty rather than by role.
Notice what these share: subagents either read without writing, or hand over a finished artefact. None of them require two agents to hold a consistent shared picture of a moving target.
Patterns that fail
- Splitting one implementation task across agents. The Flappy Bird failure. Interfaces, naming and assumptions have to agree, and prose briefs do not carry enough of them.
- Agents conversing to reach consensus. Expensive, slow, and prone to converging on a confident wrong answer. Two models agreeing is not evidence.
- Role-play org charts. A "product manager agent", an "architect agent" and an "engineer agent" mostly generate meeting minutes. Roles are a human coordination device, not a capability.
- Deep hierarchies. Every layer loses context and adds latency. Beyond two levels, debuggability collapses.
If you do split, design the interface
The failure mode is always context, so treat what crosses the boundary as the actual design work.
Briefs should be over-specified. Include the shared decisions already made — the file layout, the naming convention, the visual style, the API contract — not just the goal. Anything a subagent has to invent is something the others will invent differently.
Prefer artefacts over conversation. A subagent that writes its findings to a file gives you something reviewable, resumable and cheap to pass on. A subagent that returns 8,000 tokens of narrative gives you a context problem.
And keep results structured. Fixed fields — what was found, sources, confidence, open questions — make synthesis mechanical instead of another reasoning task for the orchestrator.
Debugging gets meaningfully harder
A single agent failure has one trace. A multi-agent failure has a tree, and the useful question is usually not which agent erred but which brief was ambiguous.
Budget for that up front: give every subagent a correlation ID, log the full brief and full result, and make the orchestrator's decomposition inspectable. If you cannot reconstruct why the lead agent split the work the way it did, you cannot fix the class of bug that matters most here.
The decision rule
- Can the work be split into parts that never need to agree? If no, use one agent.
- Is the bottleneck breadth of search rather than depth of reasoning? If no, use one agent with a bigger budget.
- Is the task worth roughly an order of magnitude more tokens? If no, use one agent.
- Can you write briefs that carry every shared assumption? If no, use one agent.
Four yeses and an orchestrator-worker design will likely beat a single agent. Anything less and you are buying coordination failure at a premium. Start with one agent, find the specific bottleneck, and split only along that seam.
Common questions
Why did Anthropic and Cognition publish opposite advice?
They were describing different tasks. Anthropic evaluated parallel research, which decomposes cleanly. Cognition was building a coding agent, where subtasks share assumptions that separate contexts cannot keep aligned.
How much more do multi-agent systems cost?
Anthropic reported roughly 15 times the token usage of a chat interaction, and found token spend explained about 80 percent of the performance variance. Much of the gain is simply spending more compute.
When is a subagent worth it for coding?
For read-only exploration. Sending a subagent to search a large codebase and return only the relevant lines keeps noise out of the main context. Splitting the implementation itself is where it breaks down.