Managing Agent Memory and Context
Agents forget, drift, and drown in their own context. How memory actually works, what fills it fastest, and patterns that keep long runs coherent.
Published 2026-09-29 · 10 min read
An agent that runs for five minutes is a different creature from one that runs for five hours. The failure mode nearly everyone hits around the second week of serious agent use looks like this: the agent starts strong, follows instructions precisely, makes real progress — and then slowly drifts. It forgets constraints stated clearly at the beginning. It redoes work it already finished. It contradicts findings it established an hour ago, with total confidence. Almost every instance of this traces back to the same root cause: memory and context management. What the agent can still see versus what has silently fallen out of its working memory determines whether a long run stays coherent or decays into expensive wandering. This guide explains how agent memory actually works, what fills it up fastest, and the practical patterns that keep long-running agents on track.
How agent memory actually works
Strip away the marketing and an AI agent's memory is straightforward: a context window, the slice of text the model can consider when producing its next action. Everything the agent knows about your task — the original instructions, the conversation so far, every tool result received — must fit inside that window simultaneously. The system prompt, the running transcript, and the latest tool output all compete for the same limited space. When the window fills, something gives: most frameworks silently drop the oldest content, and some summarize it first. Either way, the agent has no reliable sense of what it has forgotten. It will not warn you that the constraint from step two is gone; it will simply stop following it. Understanding this one mechanism explains most complaints of the form 'the agent got stupid halfway through.'
What fills up the context fastest
- Verbose tool outputs. A single unfiltered API response, a full HTML page scraped for one fact, or a directory listing of thousands of files can each consume a large share of the window in one shot.
- Retry loops. Every failed attempt re-sends the error, the reasoning about it, and the next attempt. Three retries of a bad approach can cost more context than the productive portion of the entire run.
- Retrieved documents. Pulling whole PDFs or long web pages just in case is the fastest way to drown an agent. Retrieve narrowly or not at all.
- Conversation history. Clarifications, corrections, and back-and-forth accumulate before the real work starts, so the work begins with a half-full window.
- Bloated system prompts. Every edge case, every never-do-X clause, every formatting rule sits in the window for the entire run. Prompts accrete rules over time; audit them.
- Re-reading the same sources. Agents re-open files and re-fetch pages they already processed, paying the full context cost a second time for zero new information.
- Intermediate reasoning. Long chains of thought are useful, but an agent that thinks out loud for paragraphs before each action burns window on process instead of progress.
Keep a written state file
The single highest-leverage memory pattern is embarrassingly low-tech: make the agent maintain a markdown notes file as it works. Call it state.md, progress.md, or notes.md — the name does not matter. What matters is the discipline: the agent writes down decisions made, facts established, files changed, and what remains to be done, updating the file at natural checkpoints. Because the file lives on disk rather than in the context window, it survives truncation. When the window fills and old transcript drops away, the agent re-reads the state file and recovers its bearings in seconds. Tell the agent explicitly to update the state file before finishing each subtask. Without that instruction most agents will not do it on their own, and the file goes stale exactly when it is needed most.
Checkpoint with summaries
For runs longer than about twenty steps, add explicit summarization checkpoints. Instruct the agent: every ten steps, or before switching subtasks, write a tight summary of what was accomplished, what was decided, and what is next — then continue from the summary rather than the raw history. A good checkpoint summary is five to ten lines: outcomes, open questions, and the immediate next action. It is not a transcript. Discard aggressively: dead ends, corrected mistakes, and exploratory tool calls that went nowhere should not survive into the summary. Some frameworks summarize automatically; if yours does, read what it produces for a few runs before trusting it, because auto-summaries have a habit of preserving trivia while dropping the one decision that mattered.
Trim tool outputs aggressively
- Ask for exactly what you need. Returning only the status code and the error field beats dumping a whole response and hoping the agent finds the relevant part.
- Filter at the source. Pipe through search or truncation before output reaches the agent rather than handing over megabytes and asking it to skim.
- Cap output length in tool wrappers. A hard cap forces concise results and surfaces the tools that chronically over-return.
- Prefer structured data sources over scraping. One clean data feed with the fields you need replaces pages of markup the agent must parse and discard.
- Summarize large outputs immediately. If a tool must return something big, have the agent distill the essential facts into the state file right away, then move on.
- Page through results instead of fetching everything. Several targeted queries cost less context than one give-me-everything call.
Separate memory by lifetime
Not all memory belongs in the same place, because not all of it has the same lifespan. Think in three tiers. Run-scoped memory — the current task's goals, findings, and next steps — belongs in the state file and the active context. Project-scoped memory — coding conventions, where credentials live, how deployments work — belongs in persistent project files the agent reads at startup, not in every prompt. Permanent memory — your preferences, things you always want done a certain way — belongs in the agent's long-term configuration. Mixing tiers is how windows fill with irrelevance: project conventions pasted into every run's prompt, or run-specific findings saved as if they were permanent rules. When the agent starts remembering things that applied to one old task, that is tier confusion, and it is worth cleaning up.
A starter setup you can copy
For one concrete recipe: create a state file template with four headings — Goal, Established facts, Decisions made, Remaining work. Add two lines to the agent's instructions: update the state file before finishing each subtask, and when context feels long, summarize progress into the state file and continue from the summary. Cap tool outputs at a few thousand characters. Keep the system prompt under one page by moving edge cases into project docs the agent reads on demand. Review the state file yourself after the first few long runs; you will quickly see whether the agent writes useful notes or bureaucratic filler, and you can tighten the instruction accordingly. This setup takes ten minutes and prevents the most common failure mode in long agent runs.
Memory management is the least glamorous part of working with agents and the part that most determines whether they are dependable. Nobody demos a state file. But the gap between an agent that impresses for ten minutes and one you trust with real work is almost entirely here: in what it remembers, what it forgets, and whether the forgetting is managed or accidental. Get the plumbing right and everything downstream — reliability, cost, quality — gets easier.
Triage when the context is already full
Sometimes you inherit a run that is already drowning — the agent repeats questions you answered, re-reads files it already read, or proposes steps it already tried. That is the signature of a full context: the model is working from a compressed, lossy view of its own history and filling gaps with guesses. The fix is triage, not continuation. First, stop the run — every additional step taken in a degraded state adds noise that makes recovery harder. Second, ask the agent to write down everything it currently believes to be true: the goal, what has been established, what remains. Third, read that summary yourself and correct it — this is the highest-leverage five minutes in the whole process, because everything downstream inherits your corrections. Fourth, start a fresh run with the corrected summary as the starting context and the original goal restated. Fifth, put the guardrails in place that would have prevented the situation: a state file, output caps, and a summarization checkpoint. Triage feels like losing progress; it is actually the fastest route back to reliable work.
- The agent asks for information it was already given — file contents, decisions, constraints from earlier in the run.
- It re-runs commands or re-reads files to be sure, duplicating completed work.
- Summaries of progress become vague or subtly wrong about what was actually decided.
- It proposes steps that contradict earlier findings without acknowledging the conflict.
- Tool outputs get skimmed instead of used — the agent acts as if it never saw results it received.
The handoff problem: memory across sessions
Most real work does not fit in one session. You stop for the day, the conversation compacts, or a different agent picks up the task — and the new session starts with a fraction of the old one's understanding. Handoff notes are the bridge, and they need a specific shape to work. A good handoff contains the goal in one sentence, the current state in three to five bullets, every decision made and why (this is what gets lost most often), the known dead ends so the next session does not re-explore them, and the precise next step. Write handoffs for a competent stranger, because that is effectively who reads them — a future session with no memory of your collaboration. The common failure is writing handoffs as narratives (we tried X, then Y happened) instead of state (X is decided; Y is the next step). Narratives require reconstruction; state can be acted on immediately. If your workflow regularly spans sessions, make the handoff note a required deliverable of every session, not an afterthought.
Measuring whether your memory setup actually works
Memory setups feel productive to build and are rarely tested. Test yours with three questions after a long run. First: did the agent ever ask for something it had already been told? Each repeat is a memory failure you can trace to a specific missing note. Second: after a summarization checkpoint, did the quality of subsequent work stay level or improve? If it dropped, your summaries are losing critical detail — usually decisions and their reasons. Third: could a new session pick up from your state file alone? Try it once: hand the state file to a fresh session and see how long before it asks a question the file should have answered. Run this audit monthly on your most-used workflows. The fixes are usually small — a missing heading in the state template, a summarization instruction that drops rationale — but you only find them by checking.
Problems that look like memory failures but are not
Not every case of the agent forgetting is a memory problem. Sometimes the instruction was never clear enough to be remembered — vague goals produce vague recall, and no state file fixes a goal the agent never understood. Sometimes the context is full of the wrong things: three competing instruction documents, verbose tool definitions, or a system prompt that restates the obvious at length. The agent remembers fine; it is drowning in noise. Sometimes the real issue is too many tools: with forty tools available, the agent mis-selects and the failure looks like forgetting how to do the task. And sometimes it is simply a weak model for the task's complexity — no memory architecture compensates for a model that cannot hold the reasoning chain. Before rebuilding your memory setup, check these four. The cheapest fix is usually deleting irrelevant context, not adding more memory machinery.
Keep reading
A practical look at where AI agents genuinely help in 2026 and where they still fall short, so you can delegate with confidence.
How to Write Instructions AI Agents ActuallyClear, practical techniques for writing agent instructions that get followed — with examples of vague versus precise wording.
AI Agent vs. Chatbot vs. Script: What's theAgents, chatbots, and scripts get lumped together. Here is what each one actually is, and when to use which.