Why an AI Agent Loses Track During a Long Task
Context window room is not the same as attention. Here is why long agent sessions drift and what fixes it.
Why does an AI agent lose track during a long task it was handling correctly an hour earlier? An agent that was reading files correctly and making sensible edits starts re-reading the same file for the third time, forgets a constraint you gave it at the start, or reintroduces a bug it already fixed. Nothing crashed. The model did not get replaced mid session. What changed is how much of the conversation is still doing useful work by the time it reaches that point.
The Context Window Fills With Its Own Work
Every tool call, every file read, every command output an agent runs during a long task gets appended to its context, the same window covered in what a 4 million token context window means. A short task barely touches that budget. A long one, dozens of file reads, several failed attempts, tool output pasted back in, can burn through a large fraction of it well before the task is done, and everything already in the window competes for the model's attention on every subsequent step.
This is different from running out of context outright. A model can still have room left and still perform worse, because the signal that mattered, your original instruction, a constraint stated once near the start, a decision made ten tool calls ago, is now a small fraction of a much larger window dominated by verbose command output and repeated file contents.
Why Coding Agents Hit This Especially Hard
They read entire files, not summaries, and a handful of medium files can be a meaningful chunk of a working context budget on their own.
Failed attempts do not disappear. A wrong edit, the error it caused, and the fix all stay in the transcript, which is exactly why the model sometimes repeats a mistake it technically already saw fail once.
Tool output is often noisy: full stack traces, full test runner logs, full linter output, when only the last few lines usually mattered.
Some of this overlaps with why an agent reads more files than seems necessary in the first place, covered in why does an AI coding agent read so many files, but the failure mode here is different: the agent read the right things, they are just getting crowded out later in the same session.
What Actually Helps
Break large tasks into smaller sessions with a clear handoff. Summarize what was decided and why at the end of one session, and start the next one from that summary instead of the full transcript.
Restate hard constraints close to the point where they matter, not just once at the very start of a long session.
Ask the agent to summarize its own plan and progress periodically. A summary the model writes itself is a cheap way to compress the useful signal back down before it gets buried.
Trim tool output before it goes back into context where your harness allows it: last 20 lines of a log instead of the full 400, a diff instead of a full file re-read.
None of this is about the model getting worse. It is about the ratio of relevant instruction to accumulated noise getting worse as a session runs longer, which is a property of the conversation, not the model's ability.
A Worked Example
Say you start a session telling an agent to never touch the payments module without asking first, then spend the next two hours on an unrelated refactor across forty files. By the time a change legitimately touches something adjacent to payments, that constraint is one sentence sitting far back in a transcript dominated by forty files of unrelated diffs and command output. The agent is not choosing to ignore the rule. The rule is now a small, weakly weighted signal competing against a much larger, much more recent body of context describing a completely different task. Restating the constraint right before the payments-adjacent change, rather than trusting the original instruction to still carry weight two hours later, is what actually prevents the mistake.
The same pattern explains why an agent that handled a tricky edge case correctly early in a session sometimes handles the identical case wrong later on. The correct handling is not gone, it is buried under everything that happened afterward, and buried instructions do not reliably win against a model's more immediate read of what the conversation seems to be asking for right now.
FAQ
Is this the same thing as running out of context window?
No. An agent can still have context room left and still perform worse on a long task, because the useful instructions are now a small fraction of a much larger window full of tool output and repeated file reads.
Does switching to a model with a bigger context window fix this?
It raises the ceiling before the window fills, but it does not fix the underlying issue: more room to accumulate noise is not the same as the model weighing old instructions correctly once that noise builds up.
Should I just start a fresh session for every task?
For long or multi-stage tasks, breaking at natural checkpoints with a written summary handoff works better than either one giant session or restarting from nothing each time.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


