Agent Memory: Managing Context Before It Manages You
AI Agents

Agent Memory: Managing Context Before It Manages You

Long agent sessions fail because context fills with noise. Here are the practical strategies for deciding what an agent should remember and what to drop.

Models are stateless. Every apparent memory an agent has is something your harness re-sent. That makes memory a design problem: you decide what survives into the next turn, and that decision determines whether a long session degrades or stays sharp.

Why long sessions get worse

Three things compound. Context fills with tool output that mattered twenty turns ago. Relevant instructions drift into the middle, where attention is weakest. And every turn costs more, because it carries everything before it.

The symptom is familiar: an agent that was precise early becomes vague, forgets a constraint you stated at the start, or re-reads a file it already read.

Four kinds of memory

  • Working context — the current conversation. Fast, expensive, finite.
  • Summarised history — compacted older turns. Cheap, lossy.
  • Retrieved memory — facts fetched on demand from a store. Scales indefinitely, costs a lookup.
  • Durable artifacts — notes the agent writes to disk. Survives restarts, reviewable by humans.

Most production agents use all four. The skill is routing each fact to the right one.

Compaction is the workhorse

When context approaches a threshold, replace old turns with a summary. What to preserve:

  1. The original objective, verbatim. This is the thing most often lost, and losing it is fatal.
  2. Decisions made and why.
  3. Constraints discovered along the way.
  4. What has already been tried and failed.

What to drop: full file contents, verbose tool output, superseded reasoning. Keep the conclusion, discard the transcript that produced it.

That fourth item matters more than it looks. Without a record of failed approaches, a compacted agent cheerfully retries the thing that already did not work.

Write things down

The most robust memory is a file. An agent that maintains a scratch document — current plan, findings, open questions — gets several benefits at once: it survives compaction, it survives a crash, and a human can read it.

This also turns memory into something reviewable. When an agent goes wrong, the notes usually show exactly where its model of the problem diverged from reality.

Do not remember everything

The instinct with retrieval-backed memory is to store every interaction. This degrades quickly: the store fills with stale, contradictory facts, and retrieval starts surfacing things that were true three sessions ago.

Prefer storing decisions and stable facts over transcripts. Attach a timestamp and prefer recent entries on conflict. Periodically prune. A small accurate memory beats a large stale one.

Practical thresholds

  • Compact at roughly 60–70% of the window, not at 95%. Late compaction means you are already operating in the degraded zone.
  • Truncate individual tool results hard — a few thousand tokens each, with an explicit note that output was cut.
  • Re-state the objective after each compaction, at the end of context where attention is strong.
  • Cap total turns. A session that has run eighty turns is usually lost rather than nearly finished.

Not everything belongs in memory

A useful filter before persisting anything: will this still be true next week, and would an agent be wrong without it?

Stable facts pass — where the auth logic lives, why a library was chosen, which approach failed and why. Transient state fails: the contents of a file at a moment in time, an intermediate calculation, the output of a command that will be re-run anyway.

Persisting transient state is how memory stores become actively harmful. The agent retrieves a file's contents from three sessions ago, treats it as current, and reasons from a stale premise with total confidence.

Handling the restart case

Long-running agents get interrupted — a crash, a timeout, a deploy. What survives determines whether resumption means continuing or starting over.

This is the strongest argument for durable artifacts. An agent whose plan and findings live in a file can be restarted with a single instruction to read it and continue. One whose state existed only in the conversation has to rediscover everything, at full token cost.

Memory across sessions

Session-scoped memory is the easy case. Cross-session memory — an agent that remembers your preferences and your codebase's quirks weeks later — is where most designs come unstuck, because relevance decays and nothing prunes it.

A workable pattern is to keep cross-session memory small, explicit and human-editable. A checked-in conventions file beats an opaque vector store for most teams: you can read it, correct it, and review changes to it in a pull request.

Test it deliberately

Memory bugs hide in short tests. Run a task long enough to trigger compaction, then check whether the agent still respects a constraint you set at the very beginning. That single test finds most memory design errors, and it is the one almost nobody runs.

Common questions

How much context should I use before compacting?

Around 60 to 70 percent of the window. Waiting until it is nearly full means you spend the final turns in the range where attention and accuracy are already degraded.

Should agents store every conversation in a vector database?

No. Stores full of raw transcripts surface stale and contradictory facts. Save decisions and durable facts, timestamp them, prefer recent entries, and prune regularly.

What is the single most important thing to preserve when compacting?

The original objective, word for word, plus the list of approaches already tried and failed. Losing either causes the agent to drift or repeat itself.

Similar articles

Agent Cost Control: Patterns That Bound the Worst Case
AI Agents
AI Agents·9 min read

Agent Cost Control: Patterns That Bound the Worst Case

Agent spend is not a per-call problem, it is a per-loop problem. Budgets, caching, tiering and circuit breakers that cap what a bad run can cost you.

Read
Agent Failure Modes: A Taxonomy Worth Memorising
AI Agents
AI Agents·9 min read

Agent Failure Modes: A Taxonomy Worth Memorising

Agents fail in about eight recognisable ways, and each one needs a different fix. A field guide to spotting them from a trace and knowing what to change.

Read
Agent Observability: Tracing a Loop You Cannot Reproduce
AI Agents
AI Agents·9 min read

Agent Observability: Tracing a Loop You Cannot Reproduce

Agent failures are rarely reproducible, so logs are not enough. What to record per step, how to span a tool loop, and which metrics predict a bad run.

Read