When Not to Use an Agent
AI Agents

When Not to Use an Agent

Most tasks handed to agents are better served by a fixed workflow or a single model call. A decision rule, the failure modes, and what to build instead.

Agents are the most interesting thing you can build with a language model, which is why so many systems contain one that should not. The interesting architecture and the correct architecture are frequently different.

Anthropic makes this point in its own guidance on building effective agents: find the simplest solution possible, and increase complexity only when it demonstrably improves outcomes — which may mean not building an agentic system at all. Agentic systems trade latency and cost for task performance, and that trade is not always worth making.

The distinction that matters: who decides the control flow

A workflow orchestrates model calls through code paths you wrote. You decide what runs, in what order, with what branching. The model fills in the intelligent parts.

An agent lets the model decide the control flow. It chooses which tool to call, when, and when to stop.

Nearly everything about a system follows from that choice. Workflows are testable, cheap, predictable and boring. Agents are flexible, expensive, non-deterministic and interesting. You want the boring one unless you specifically need the flexibility.

Reliability compounds against you

This is the argument that decides most cases, and it is arithmetic rather than opinion.

An agent that picks the right action 95% of the time is not a 95% reliable agent. Over twenty dependent steps it is about 36% reliable, because errors compound. Push per-step accuracy to 99% and twenty steps still only gets you to about 82%.

Anything you can do to reduce the number of model-made decisions improves reliability more than any prompt change will. A workflow with three model calls in fixed positions has three chances to go wrong. An agent solving the same problem has however many it takes.

Six cases where an agent is the wrong tool

1. The steps are always the same. If you can write the sequence down and it does not change, write it down. Extract, then classify, then format, then store. Giving that to an agent means paying a model to rediscover your own procedure on every request, and occasionally getting it wrong.

2. The task is one-shot. Summarise this, translate that, classify these. There is nothing to find out, so the loop adds round trips and latency for no information gain. One call, done.

3. Latency has a hard ceiling. Anything in a user-facing request path where you have a budget in the low hundreds of milliseconds. Agents are unbounded by construction — you cannot promise a response time for a loop whose length the model chooses.

4. Errors are expensive and unreviewable. Payments, deletions, communications sent to customers, anything touching production infrastructure. If a human is going to check the output anyway, a workflow that produces a reviewable artefact is better than an agent that acts and reports.

5. You need the same answer every time. Compliance, pricing, entitlement decisions, anything auditable. Agents are non-deterministic in path as well as in wording. Two runs of the same input can take different routes to different answers, and "the model chose to" is not an audit trail anyone accepts.

6. Volume is high and margin is thin. A single agent run can consume several times the tokens of a chat turn, and multi-agent fan-out multiplies that again. Anthropic reported roughly fifteen times chat token usage for its multi-agent research system. At scale, that difference is your product margin.

The signals that you do want an agent

To be fair to the other side, there is a real case, and it is narrow but genuine:

  • The next step depends on what you find. Debugging is the canonical example. You cannot write the sequence in advance because it branches on the stack trace.
  • The input space is genuinely open. Enumerating the cases in code would mean enumerating an unbounded set.
  • Verification is cheap and automatic. Tests, compilers, schema validation. When the agent can check its own work, the compounding error problem partly reverses — a wrong step gets caught and retried rather than propagating.
  • The value per task justifies the variance. Deep research, complex refactors, investigations. High value absorbs both the cost multiplier and the occasional failed run.

Notice that three of those describe exploration and one describes economics. If your task is not exploratory, the case for an agent is weak regardless of how well it would demo.

Build in this order

  1. Single model call. Try it first. Surprisingly many "agent" requirements are one well-specified prompt with a structured output schema.
  2. Chained calls. Fixed sequence, each step feeding the next. Add a validation gate between steps and you can retry just the step that failed.
  3. Routed calls. Classify the input, dispatch to a specialised prompt. Deterministic branching, model-driven classification.
  4. Workflow with tools. Your code decides which tools run; the model interprets the results. Most of the benefit of an agent with none of the unbounded loop.
  5. Agent. Only when the four steps above genuinely cannot express the task.

Each level up costs more, fails in more ways and takes longer to debug. The discipline is to move up only when you have evidence the level below is insufficient — not when you suspect it might be.

A decision rule you can apply in a minute

Ask: can I write down the steps before seeing the input?

If yes, build a workflow. If no, ask a second question: does the next step depend on information only obtainable by acting? If yes, you want an agent. If no — if the steps vary but are enumerable — you want routing.

The honest summary is that agents are the right answer for a smaller share of tasks than current enthusiasm suggests, and the wrong answer in ways that are expensive to discover late. When they are right, they are transformative. Getting there by elimination is cheaper than getting there by rewrite.

Common questions

What is the difference between an AI workflow and an AI agent?

In a workflow, your code decides the control flow and the model fills in the intelligent parts. In an agent, the model decides the control flow — which tools to call, in what order, and when to stop. Workflows are predictable and testable; agents are flexible and non-deterministic.

Why do agents become unreliable over many steps?

Errors compound. An agent that is 95% accurate per action is only about 36% reliable across twenty dependent actions. Reducing the number of model-made decisions improves reliability more than almost any prompt change.

When is an agent clearly the right choice?

When the next step genuinely depends on what the previous step found, and when the agent can verify its own work cheaply — running tests, compiling, validating a schema. Exploratory debugging and multi-file refactoring are the clearest cases.

Similar articles

AI vs AI Agent: What Actually Changes When You Add a Loop
AI Agents
AI Agents·8 min read

AI vs AI Agent: What Actually Changes When You Add a Loop

A model answers. An agent decides, acts, and checks its work. The difference is a loop and a set of tools — here is exactly what that means in code.

Read
Agent Cost Control: Patterns That Bound the Worst Case
AI Agents
AI Agents·9 min read

Agent Cost Control: Patterns That Bound the Worst Case

Agent spend is not a per-call problem, it is a per-loop problem. Budgets, caching, tiering and circuit breakers that cap what a bad run can cost you.

Read
Agent Failure Modes: A Taxonomy Worth Memorising
AI Agents
AI Agents·9 min read

Agent Failure Modes: A Taxonomy Worth Memorising

Agents fail in about eight recognisable ways, and each one needs a different fix. A field guide to spotting them from a trace and knowing what to change.

Read