Agent Planning Strategies: From ReAct to Explicit Plans
Reactive loops drift on long tasks and rigid plans break on contact with reality. Here are the planning strategies that work, and how to choose between them.
An agent without a plan does the first plausible thing, then the next, and forty steps later has solved a problem nobody asked about. An agent with a rigid plan follows step three into a wall because the plan was written before step two revealed the wall.
Planning strategy is the choice between those failure modes. There are a handful of well-understood options and they suit different tasks.
Reactive: think, act, observe
The ReAct pattern interleaves reasoning and action in a single loop. The model produces a thought, calls a tool, sees the result, and repeats. There is no plan document — the plan exists implicitly in the sequence of decisions.
This is the default for most coding agents and it works well for exploratory work, because every step is informed by the last observation. It adapts continuously and has almost no upfront cost.
Its weakness is horizon. Because thought and action fire in the same turn, there is no checkpoint before committing, and nothing holds the overall objective in view. On long tasks a reactive agent drifts: it chases the most recent interesting detail, loses the original goal, and rediscovers dead ends it already visited.
Plan then execute
Generate an explicit ordered plan first, then work through it. The planning call sees the whole task; the execution calls each see one step.
Three real benefits. The plan is a human review point before anything expensive happens. Steps can be executed by cheaper models or in parallel where they are independent. And the objective survives compaction, because it lives in the plan rather than in the drifting conversation.
The cost is brittleness. A plan written before execution encodes assumptions about a world the agent has not inspected yet. Step four says "update the config in settings.json" and there is no such file. Without a mechanism to revise, the agent either improvises silently or fails.
Plan with replanning
The practical middle. Generate a plan, execute a step, then check whether the plan still holds. Revise when it does not.
The design question is what triggers a revision, because checking after every step costs a full model call and usually finds nothing. Useful triggers:
- A step failed, or produced a result contradicting an assumption in the plan.
- A discovery invalidates a later step — a missing file, a different framework, an API that does not exist.
- The step budget is more than half consumed with less than half the plan done.
- The same step has been attempted twice.
Replan by editing the remaining steps, not by regenerating from scratch. A full regeneration loses the record of what has already been tried, which is how agents end up cycling through the same three failed approaches.
Search over plans
When a single line of reasoning is not enough, you can explore several. Tree of Thoughts generalises chain-of-thought into a search over intermediate steps with self-evaluation and backtracking. On the Game of 24 arithmetic puzzle, the paper reported 74 percent success against 4 percent for GPT-4 with chain-of-thought prompting.
That is a real result on a task with a cheap, exact verifier. Note the condition. Search only pays when you can evaluate a partial path — a test suite, a type checker, a puzzle constraint. Without a verifier, the model scores its own branches and confidently prunes the good one.
Search also multiplies cost by the branching factor. For most production agents it is the wrong tool; for constrained problems with automatic checking it can be transformative.
Reflection between attempts
Reflexion adds a different axis: after a failed attempt, the agent writes a short analysis of what went wrong and carries that text into the next attempt. Learning across trials, in natural language, without touching weights.
This is cheap and effective when failure is detectable — tests fail, the build breaks, the API returns an error. It is much weaker when the agent has to judge its own success, since a model that could recognise its mistake would often have avoided it.
Keep reflections short and concrete. "The regex failed because filenames contain dots" is useful. "I should be more careful" is filler that costs tokens and changes nothing.
Generating versus verifying
The most useful framing from the planning research literature is that language models are far better at proposing plans than at guaranteeing them. Work around PlanBench has repeatedly found autoregressive generation unreliable as a standalone planner, and the LLM-Modulo line of work argues for pairing a model generator with an external critic that checks the candidate.
The engineering translation is straightforward: never let the model be the only judge of whether the plan is sound. Wherever a real verifier exists, use it. Type checkers, tests, schema validation, a dry-run mode, a linter. A model proposing and a machine checking beats a model doing both.
Make the plan an artefact
Write the plan to a file rather than keeping it in the conversation. This is a small change with outsized effects.
It survives compaction and restarts. A human can read and edit it mid-run. Marking steps done gives progress that is visible without parsing a transcript. And when the agent goes wrong, the diff between the original plan and the current one usually shows exactly where its model of the problem diverged.
The same file is the natural home for the record of failed approaches — the single most valuable thing to preserve, and the thing most often lost.
Choosing
- Short, exploratory, unclear scope — reactive. Planning overhead exceeds the benefit under roughly five steps.
- Long, multi-file, known shape — plan with replanning, plan in a file, human review before execution.
- Verifiable and constrained — add search or repeated attempts, gated by the verifier.
- Repeated failures on a known task — add reflection carried between attempts.
- Irreversible actions anywhere in scope — always plan first, and gate the destructive steps on human approval.
Default to reactive and add structure only where you can point at the failure it fixes. Every planning layer is more tokens, more latency and more surface area, and an agent over-planning a two-step task is its own kind of failure.
Common questions
Should an agent always plan before acting?
No. Under about five steps the planning call costs more than it saves and the plan is usually discarded anyway. Planning earns its keep on long, multi-step tasks with irreversible actions.
Why do plans fall apart mid-execution?
They encode assumptions made before the agent inspected the environment. The fix is a replanning trigger on failed steps and contradicted assumptions, editing the remaining steps rather than regenerating.
Is tree search worth the cost?
Only with a cheap, reliable verifier for partial solutions. Tree of Thoughts reported 74 percent on Game of 24 against 4 percent for chain-of-thought, but that task has an exact checker. Without one the model prunes its own good branches.