Long-Running Agents: Durability, Compaction and Restarts
An agent running for hours hits limits a chat never does: context exhaustion, process restarts, stale state. The patterns that keep long jobs alive.
An agent that runs for thirty seconds and an agent that runs for six hours are not the same system with a different timeout. Past roughly the first hour, three things start failing that never come up in a chat: the context fills, the process dies, and the world the agent reasoned about changes underneath it.
None of those are model problems. They are runtime problems, and they are solved in your code.
Context is a budget, and it runs out
Every tool result stays in the transcript. A long run accumulates file contents, command output, stack traces and its own reasoning until the window is full — and because models are stateless, you resend all of it on every step.
Three strategies, in increasing order of how much work they take:
Prune. Drop what is provably dead. Tool results that were superseded, file reads for files since edited, failed command output already acted on. Cheap and safe, and usually the first thing worth doing.
Compact. Summarise older turns into a shorter block and keep the recent window verbatim. The subtlety is what survives: decisions and constraints must survive, narrative does not need to. A summary that records "tried approach A, failed because of the auth middleware" is worth ten summaries that say "explored several options".
Externalise. Move state out of the context entirely, into a file the agent reads and writes. A structured working document — goal, constraints, decisions, open questions, next step — lets you reset context completely and rehydrate from the document. This is the only approach that survives arbitrarily long runs, because the memory lives outside the window.
At day-plus durations, summarisation alone stops being enough. Plan for full resets driven by an explicit handoff artefact rather than assuming ever-better compaction.
The process will die
Deploys, OOM kills, spot instance reclamation, network partitions. Over six hours, something interrupts. The question is whether that costs you a step or the entire run.
Two mechanisms dominate, and they are worth naming precisely:
Journal and replay. Append every completed step and its result to a durable log. On restart, replay the journal to rebuild in-memory state without re-executing anything, then continue from the first unfinished step. This is how durable execution engines work, and several agent frameworks — LangGraph, Pydantic AI, the OpenAI Agents SDK among them — now expose durability as a first-class feature.
Checkpointing. Persist the full agent state after each step. Simpler to reason about, larger to store, and it makes forking trivial: you can branch a run from step 40 and try a different approach without redoing the first 39 steps.
Either way, the invariant is the same: persist at completed step boundaries, and never re-execute a side effect on recovery.
Idempotency is the hard part
Replay only works if replaying is safe. Rebuilding state from a journal must not re-send the email, re-charge the card, re-open the pull request or re-post the message.
The practical rules:
- Treat tool executions as recorded facts. On replay, read the recorded result rather than calling the tool again.
- Give every side-effecting call an idempotency key derived from the run identifier and step index, so a genuine retry is deduplicated upstream.
- Write the journal entry before the effect and update it after, so a crash mid-call leaves evidence that something might have happened.
- Separate read tools from write tools in your registry. Reads are free to retry; writes are not.
A crash between "sent the request" and "recorded the result" is the ambiguous case, and there is no clever trick that removes it. You either make the operation idempotent upstream or you pause for a human.
Long runs drift
The most insidious failure is not a crash. It is an agent that at hour four is confidently working from an understanding formed at hour one, on a codebase that has since changed — sometimes changed by the agent itself.
Counter it by re-grounding on a schedule. Periodically discard the accumulated belief and re-read the current state: run the tests again, re-read the file rather than trusting the earlier read, restate the objective from the original brief rather than from the summary of the summary.
It is worth building a cheap verification step into the loop for exactly this. An agent that runs the test suite every ten steps notices breakage within ten steps. One that runs it at the end notices after six hours of work built on it.
Bound the worst case explicitly
Long-running does not mean unbounded. Every long agent needs hard ceilings enforced by the harness, not requested in the prompt:
- A step cap and a wall-clock cap.
- A token or cost budget for the run, checked before each model call.
- A no-progress detector: if the last N steps produced no state change, stop.
- A repeated-action detector: identical tool call with identical arguments twice is a stall signal.
Hitting a ceiling should not discard the work. Persist the state, emit the handoff document, and surface it for a human to resume or redirect. A run that stops cleanly at 80% with a written summary is far more useful than one that grinds to the cap and returns nothing.
This is also where the billing model starts to bite. Metered per-token pricing makes long, exploratory runs a spending decision on every step, which pushes teams to cut them short for reasons that have nothing to do with the task. Flat-rate access removes that particular pressure — it does not remove the need for step caps, which exist to catch stalls, not to protect the invoice.
Humans belong in the loop, at boundaries
The natural checkpoints are phase transitions: plan approved, migration about to run, pull request about to open. Pausing there is cheap, because state is already persisted and resuming is a normal operation.
Design the pause as a first-class state, not an exception. An agent that can serialise itself, wait a day for approval, and resume exactly where it stopped is a much better fit for real work than one that must hold a connection open to stay alive.
Checklist
- Externalise long-term state into a document; do not rely on compaction alone.
- Persist at completed step boundaries and replay without re-executing effects.
- Give every write an idempotency key.
- Re-ground periodically instead of trusting stale reads.
- Enforce step, time and cost ceilings in code.
- Make pause-and-resume a supported state, including across deploys.
Common questions
What breaks first when an agent runs for hours?
Context. Every tool result accumulates in the transcript and gets resent on each step, so the window fills and the input cost climbs. Process crashes and stale assumptions come next, but context exhaustion is almost always the first wall.
Is summarising the context enough for very long runs?
Not on its own. Repeated summarisation loses constraints and decisions a little at a time. For day-scale work, keep durable state in an external document the agent reads and writes, so you can reset context entirely and rehydrate from it.
How do I resume an agent after a crash without repeating side effects?
Journal each completed step with its result and replay from the journal rather than re-calling tools. Give every side-effecting call an idempotency key derived from the run and step, and write the journal entry before performing the effect.