Agent Error Recovery: Patterns That Keep the Loop Alive
AI Agents

Agent Error Recovery: Patterns That Keep the Loop Alive

Errors are the normal case in an agent loop, not the exception. These are the recovery patterns that separate agents that finish from agents that spiral.

Watch enough agent traces and a pattern emerges. The agents that finish are not the ones that make fewer mistakes — they are the ones that notice a mistake and change approach. The agents that burn a hundred steps are usually repeating one failed call with cosmetic variations.

Recovery is therefore not error handling bolted on at the end. It is the main behaviour you are designing for, because in any real environment things fail constantly: files move, APIs rate-limit, tests break, assumptions turn out to be wrong.

Make errors legible

The model can only recover from what it can read. A raw stack trace, an HTTP 500 with an empty body, or a silent empty result all give it nothing to work with, so it guesses.

Return errors as instructive text:

Bad:  Error: ENOENT
Good: File not found: src/ap.ts
      Similar files: src/app.ts, src/api.ts
      Use list_directory to see what exists.

Three components make an error usable: what failed, why, and what to try next. That third part is the one people leave out, and it is the one that turns a loop into a recovery.

Errors should also be returned as normal tool results rather than thrown as exceptions that abort the run. An exception ends the episode; a tool result keeps the agent in a position to adapt.

Classify before you retry

Blind retries are the most common recovery bug. Three categories, three different responses.

  • Transient. Rate limits, timeouts, 5xx, network blips. Retry with exponential backoff and jitter, in your harness, without involving the model at all. The model should never see a 429.
  • Deterministic. Malformed arguments, a schema violation, a file that does not exist. Retrying identically will fail identically. Return the specific validation error and let the model construct a different call.
  • Semantic. The call succeeded and the result was wrong or unhelpful. No retry fixes this. The agent needs a different approach, which usually means it needs to be told that the current approach has now failed twice.

Handling transient errors below the model layer is a large and cheap win. It removes noise from context, avoids wasting a model call on a wait, and stops the agent from concluding that a tool is broken when it was merely busy.

Detect loops explicitly

Models are strongly biased toward continuing. Left alone, an agent will re-issue nearly identical calls indefinitely because continuing is more plausible than stopping.

Track a normalised signature of each tool call — name plus canonicalised arguments — and count repeats. On the second identical failing call, stop returning the same error and return an intervention instead:

This exact call has now failed twice with the same error.
Do not repeat it. Either use a different tool, change the
arguments materially, or call ask_human to report the blocker.

Naming the loop in the transcript is remarkably effective, because the model can then see the pattern it was blind to. Pair it with budgets: a hard step cap, a token cap, and a wall-clock cap, each producing a graceful summary rather than a silent halt.

Idempotency and compensation

Retries assume operations are safe to repeat. Many are not.

Give every write tool an idempotency key derived from the operation, and have your handler return the original result on a duplicate rather than performing the action twice. This is standard practice for payment APIs and it applies verbatim to agents, which retry far more enthusiastically than humans.

For multi-step operations that can fail partway, record a compensating action for each completed step — the branch to delete, the record to revert, the file to restore. Recovery then means unwinding to a known state instead of leaving half a change behind.

Half-completed work is worse than no work, because the next attempt starts from a state neither the agent nor the environment expects.

Checkpoint so a failure is not a restart

Long runs get interrupted: a crash, a deploy, a rate limit that outlasts your retries. What survives determines whether resumption means continuing or starting over at full token cost.

Persist the objective, the plan with steps marked done, findings so far, and the list of approaches already tried and failed. A file is enough. Resumption becomes a single instruction to read it and continue.

That failed-approaches list is the item most often omitted and most expensive to lose. Without it, a resumed agent cheerfully re-attempts the thing that already did not work, usually twice.

Reflect, briefly

After a genuine failure, one short reflection step often changes the trajectory: what was assumed, what was actually observed, what to try differently. The Reflexion line of work formalised this as carrying textual self-critique into the next attempt, and it works best exactly where failure is objectively detectable — tests fail, the build breaks, the API rejects the payload.

Keep it bounded. One reflection per failure, a few sentences, appended to the durable notes. Reflection loops that reflect on reflections burn tokens and produce increasingly abstract self-criticism that changes no behaviour.

Escalation is a recovery path

Some failures are not the agent's to solve. Missing credentials, a genuinely ambiguous requirement, a decision with business consequences.

Give the agent an explicit way out — a tool that reports the blocker, what it tried, and what it needs — and state in the prompt that using it is a valid ending. Without that, models grind, because stopping is not something they do naturally.

A useful trigger: two distinct approaches have both failed, or a required resource is unavailable and no alternative exists. Escalating at that point costs a few hundred tokens. Not escalating costs the remaining budget and produces a confidently wrong result.

Test recovery deliberately

Clean-path tests tell you almost nothing about production, where the paths are rarely clean. Inject failures on purpose:

  • Return a 500 from a tool the agent depends on.
  • Return a plausible but wrong result — a file that exists but has different contents than expected.
  • Return an empty result where the agent expects data.
  • Fail a write halfway through a multi-step operation.
  • Kill the process mid-run and restart it from the checkpoint.

Then score recovery rate: how often the agent adapts, escalates appropriately, or fails cleanly with an accurate report. Track it next to task success. It is the single metric that best predicts whether an agent survives contact with a real environment, and it is the one almost nobody measures.

Common questions

Should the model see tool errors, or should the harness hide them?

Hide transient ones. Rate limits and timeouts should be retried below the model with backoff. Deterministic and semantic failures belong in context, phrased so the agent can act on them.

How do I stop an agent repeating the same failing call?

Hash a normalised signature of each call and count repeats. On the second identical failure, replace the error with an explicit instruction not to repeat it and to change approach or escalate.

What should survive a crash mid-run?

The objective verbatim, the plan with completed steps marked, findings so far, and the approaches already tried and failed. Write them to a file so resuming is one instruction rather than a rediscovery.

Similar articles

Idempotency in Agent Actions: Safe to Repeat by Design
AI Agents
AI Agents·9 min read

Idempotency in Agent Actions: Safe to Repeat by Design

Agents retry, resume and duplicate calls constantly. Idempotency keys, natural keys and check-then-act patterns that stop one action happening three times.

Read
Partial Failure Recovery: When Half the Agents Succeed
AI Agents
AI Agents·9 min read

Partial Failure Recovery: When Half the Agents Succeed

Six workers, four returned, two timed out. Whether to answer, retry or abort is a design decision — here is how to make it before the incident, not during.

Read
Rollback Strategies for Agents That Changed Things
AI Agents
AI Agents·9 min read

Rollback Strategies for Agents That Changed Things

An agent stopped halfway through eleven changes. Checkpoints, compensating actions, staging areas and the undo plan you should require before it acts.

Read