ReAct vs Plan-and-Execute: Two Ways to Run an Agent Loop
AI Agents

ReAct vs Plan-and-Execute: Two Ways to Run an Agent Loop

ReAct decides one step at a time. Plan-and-execute commits to a plan first. The choice changes cost, latency, recoverability and how failures look.

Every agent has to answer one architectural question: does it decide the next action after seeing the last result, or does it decide the whole sequence up front?

That is the entire difference between ReAct and plan-and-execute. It sounds like a detail. It determines your token bill, your latency profile, whether failures are recoverable, and whether you can show a user what the agent intends to do before it does it.

ReAct: think, act, observe, repeat

ReAct came out of work by researchers at Princeton and Google in 2022, and the idea is deliberately small: let the model interleave reasoning and acting rather than doing all of one and then all of the other. At each step it produces a thought, an action, and then receives an observation, which feeds the next thought.

Thought: I need to see which tests are failing.
Action: run_tests()
Observation: 2 failed - test_auth, test_session

Thought: Both touch the session module. Read it.
Action: read_file("session.py")
Observation: ...

The explicit thought is not decoration. It decomposes the task, tracks progress, and gives the model somewhere to notice that the last observation contradicted its assumption. Remove it and tool selection gets measurably worse.

Strengths. Adapts immediately to whatever comes back. Handles tasks where the second step genuinely cannot be known until the first completes — which is most debugging, most exploration, most work against a codebase you have not read.

Weaknesses. Greedy. It optimises the next action, not the trajectory, so it will happily take twelve locally reasonable steps down a path a moment of forethought would have rejected. It is also the harder shape to display to a user, because there is no plan to show — only history.

Plan-and-execute: decide the route first

The alternative separates the phases. One model call produces an ordered plan; a cheaper loop executes each step; a replanning step runs when something fails or the plan is exhausted.

Plan:
  1. Locate the retry configuration
  2. Read the calling code
  3. Change the limit to 5
  4. Run the tests
  5. If tests fail, replan

Execute step 1 -> ...
Execute step 2 -> ...

The planning call sees the whole task at once, which is precisely what ReAct never does. That is where the gain comes from: an explicit planning layer breaks through the ceiling that a purely greedy strategy hits on complex, multi-part tasks.

Strengths. Cheaper, because execution steps do not need to re-reason about the entire objective and can often run on a smaller model. Faster, because independent steps can run in parallel. Inspectable, because you can show the plan to a human before anything executes. And recoverable, because a plan step is a natural checkpoint.

Weaknesses. Plans made without information go stale on contact. Step 3 assumed a file that does not exist. Without disciplined replanning, the executor grinds through steps that stopped making sense at step 2.

The failure modes are opposites

This is the most useful way to hold the comparison in your head.

ReAct fails by wandering. It stays locally sensible and globally lost, revisits the same file three times, and loses the thread of the original objective around step fifteen. Symptoms: high step counts, repeated tool calls, an answer that addresses something adjacent to what you asked.

Plan-and-execute fails by rigidity. It follows a plan that reality invalidated, reports the plan as complete, and returns something confidently wrong. Symptoms: low step counts, clean-looking traces, output that does not match the world.

Debugging them is different too. A ReAct failure is visible in the trace as thrashing. A plan-and-execute failure looks tidy right up until you check the result.

Cost and latency, concretely

ReAct sends the accumulated transcript on every step, so input tokens grow roughly quadratically over a run. Twenty steps means twenty sequential round trips, each carrying everything before it.

Plan-and-execute pays for one expensive planning call and then a series of cheaper, shorter executions that carry the plan and the current step rather than the full history. It also unlocks parallelism: independent steps can be dispatched concurrently, which ReAct structurally cannot do because step N plus one depends on step N.

The catch is replanning. Every replan discards work and repeats the expensive call. A plan-and-execute agent that replans on every second step is more expensive than ReAct, not less. If you see that in your traces, the plans are too specific.

What most production agents actually do

A hybrid, and not because hybrids are fashionable — because the phases have different information requirements.

Plan at a coarse grain: four to six steps, each a goal rather than an action. "Find where retries are configured" is a plan step. "Open src/http/client.py at line 40" is not — that is a decision that needs information the planner does not have.

Then run each step as a small ReAct loop with its own step cap. Adaptation happens inside a step; structure comes from the plan. Replan only when a step reports that it cannot complete, and cap the number of replans so a doomed task terminates rather than cycling.

This shape also produces good ergonomics: the plan is what you show the user, the step is what you gate on approval, and the inner loop is what you cap and instrument.

Choosing

  • Short and exploratory — under about ten steps, where you cannot know step two in advance: plain ReAct. Planning overhead buys nothing.
  • Repetitive and well understood — the same shape of task every time: plan-and-execute, or a fixed workflow with no agent at all.
  • Long and multi-part: hybrid. Coarse plan, ReAct within each step.
  • Needs human approval before acting: plan-and-execute, because there is something to approve.
  • Latency-sensitive with independent subtasks: plan-and-execute, for the parallelism.

The diagnostic if you are unsure: look at your traces and count how often the agent revisits something it already looked at. Frequent revisiting means it needed a plan. Frequent replanning means the plans were too detailed for the information available when they were made.

Common questions

What is the ReAct pattern in AI agents?

A loop in which the model alternates reasoning and acting: it writes a thought, chooses an action, receives an observation, and repeats. It was introduced by researchers at Princeton and Google in 2022 and is the default shape for most tool-using agents.

Is plan-and-execute cheaper than ReAct?

Usually, because execution steps carry the plan rather than the whole transcript and can run on a smaller model. The saving disappears if the agent replans frequently, since each replan repeats the expensive planning call and discards work.

Which should I use for a coding agent?

A hybrid. Plan at a coarse grain — a handful of goal-level steps — then run each step as a small ReAct loop with its own step cap. Coding work needs adaptation within a step and structure across steps.

Similar articles

Agent Cost Control: Patterns That Bound the Worst Case
AI Agents
AI Agents·9 min read

Agent Cost Control: Patterns That Bound the Worst Case

Agent spend is not a per-call problem, it is a per-loop problem. Budgets, caching, tiering and circuit breakers that cap what a bad run can cost you.

Read
Agent Failure Modes: A Taxonomy Worth Memorising
AI Agents
AI Agents·9 min read

Agent Failure Modes: A Taxonomy Worth Memorising

Agents fail in about eight recognisable ways, and each one needs a different fix. A field guide to spotting them from a trace and knowing what to change.

Read
Agent Memory: Managing Context Before It Manages You
AI Agents
AI Agents·8 min read

Agent Memory: Managing Context Before It Manages You

Long agent sessions fail because context fills with noise. Here are the practical strategies for deciding what an agent should remember and what to drop.

Read