Dashboard

Why Your AI Coding Agent Keeps Reintroducing Bugs

A bug fix that only lives inside one conversation doesn't survive a new session or a long context window. Here is why AI coding agents re-introduce old bugs, and the regression-test habit that actually fixes it.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
17 September 20261 min read

Why Your AI Coding Agent Keeps Reintroducing Bugs

Short answer

Your agent fixed the bug once, but the fix lived only in that conversation's context, not in a test, a comment, or a linted rule. When the context window rolls over, a new session starts, or you ask for a refactor that touches the same code, the agent re-derives the "obvious" implementation from scratch, and the obvious implementation is usually the buggy one. The fix is to convert the fix into something that survives outside the conversation: a regression test, primarily.

What's actually happening

Large context windows create the illusion of persistent memory, but nothing carries between sessions unless you explicitly save it. Inside one long conversation, an agent that fixed a race condition an hour ago can still reintroduce it, for a more specific reason: attention over a long context is not uniform. Instructions and corrections near the start of a session get diluted as more tokens accumulate, a well-documented degradation, and the original bug fix can end up read as "context" rather than "constraint" by the time the agent is generating new code near the end of the session.

Across sessions the effect is total rather than partial. A fresh session has no access to the previous one's reasoning at all unless it's in the repository, so the agent regenerates code from the same training-derived defaults that produced the bug the first time. If the naive implementation of a cache eviction policy, a date parser, or a retry loop has an edge case, the agent will hit that same edge case again with fresh eyes, because nothing about the codebase itself signals that the naive version was already tried and rejected.

The fix that actually works: a failing test first

The single highest-leverage change is to never consider a bug fixed until there's a test that fails without the fix and passes with it. This does two things a comment or a commit message cannot: it runs automatically on every future change, including ones an agent makes, and CI failure is a signal an agent reliably reacts to, far more reliably than a comment it may or may not read.

  1. Reproduce the bug in a minimal test case before accepting any fix. If the agent proposes a fix without a reproduction, ask for the failing test first.

  2. Name the test after the bug, not the feature, e.g. test_cache_eviction_handles_concurrent_writes rather than test_cache. Future agents (and future you) can grep for it.

  3. Run the full test suite as part of the agent's own workflow, not just yours, so a regression shows up inside the same session that introduced it rather than surfacing in your next manual review.

  4. For bugs with no clean unit-test boundary, such as a bad prompt output or a flaky integration, write down the specific input and expected output in a checked-in file the agent can reference, even a plain markdown table.

Secondary fixes that help

  • A CHANGELOG or known-issues note in the repo root that lists fixed edge cases in plain language. Tests catch regressions mechanically, but a short human-readable note helps an agent avoid reintroducing the same design mistake in a different function.

  • Linting rules where the bug class allows it. If the bug was a specific unsafe pattern, such as string concatenation for SQL or unguarded null access, a lint rule that flags the pattern catches it before the code is even run.

  • Shorter, more frequent sessions on code that has a history of this problem, rather than one marathon session. This works against the attention-dilution effect directly by keeping the relevant fix closer to the front of the active context.

What doesn't reliably work

Repeating the instruction in a system prompt ("remember, we already fixed the race condition in the cache layer") helps inside a single session but does nothing across sessions, and even within a session it competes with everything else in context the same way any other instruction does. Comments in code help a human reviewer far more than they help an agent regenerating a function from a prompt, since a comment describes intent but doesn't structurally block the buggy version from compiling or passing CI the way a test does.

How this connects to context management generally

This is a specific case of a broader pattern: anything you want an AI coding agent to respect reliably needs to live somewhere the agent's tooling actually checks, not somewhere it merely might read. The same principle applies to scoping how much of your codebase an agent can see, to keeping agents away from your production database, and to making sure an agent can't leak your api keys in the process: a rule stated once in a prompt is a suggestion, a rule enforced by a test, a hook, or a permission boundary is a guarantee. The full set of guardrails worth having lives under AI coding tools.

FAQ

Does a bigger context window fix this?

It reduces the mid-session version of the problem somewhat, but it does not touch the cross-session version at all, since a new session starts with none of the old one's context regardless of window size. A regression test survives both.

Should I write the test myself or have the agent write it?

Have the agent write it, but review it before accepting the fix, since a test the agent wrote to match its own fix can accidentally encode the same blind spot as the fix itself. Check that the test actually fails on the old code before the fix is applied.

What if the bug isn't easily testable, like a bad tone in generated text?

Keep a small checked-in set of example inputs and the output you consider correct, and ask the agent to check new changes against that set before calling a task done. It's less rigorous than a unit test but far more durable than relying on the agent to remember.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.