Dashboard

Make an AI Coding Agent Reproduce a Bug First

A test that runs red today is the only thing separating a real fix from a plausible edit. Plus the revert check nobody does.

Steve Jefferson
Steve Jefferson
Developer Advocate
13 September 20261 min read

Make an AI Coding Agent Reproduce a Bug First

The single highest-value rule when debugging with an agent: make it produce a failing test that reproduces the bug before it is allowed to touch the implementation. Not a description of the bug, not a theory about the cause, a test that runs red today. Without that gate you cannot tell a fix from a plausible-looking edit, because the agent will confidently report success either way, and so will your test suite if nothing in it covered the broken path.

Why agents guess, and why the guess looks right

Given a bug report, a coding agent does what the training distribution rewards: it reads some files, forms a hypothesis, and changes the code to match the hypothesis. That is a reasonable process and it is often correct. The problem is the verification step. The agent runs the suite, the suite passes, and it reports the bug fixed. But the suite passed before the change too, because no test exercised the broken behaviour. Nothing in that loop ever observed the bug.

So you get three outcomes that are indistinguishable from the outside:

  • A real fix, verified by nothing.

  • A change that alters behaviour somewhere adjacent and leaves the bug intact.

  • A change that masks the symptom, typically a defensive null check, while the bad state is still produced upstream.

The third is the expensive one, because it closes the ticket and re-opens as a stranger three weeks later. Requiring a reproduction converts all three into a single observable question: does this test go from red to green? That is a different discipline from writing a good bug report for the agent, which is about the input. This is about refusing to accept the output until something demonstrates it.

How to get an AI coding agent to reproduce a bug

  1. State the observed behaviour and the expected behaviour, with the exact input. No theory about the cause.

  2. Instruct: write a test that fails because of this bug. Do not modify any non-test file yet.

  3. Run the test. Confirm it fails, and read the failure message. It must fail for the stated reason, not because of a typo or a missing fixture.

  4. Only now allow the implementation change.

  5. Run the test again. It must pass.

  6. Revert the implementation change and run the test again. It must fail again.

  7. Re-apply the fix. Run the full suite.

Step 6 is the one everybody skips and the one that catches the real failure mode. A test written after an agent has already formed a theory has a habit of encoding the theory rather than the bug, and such a test frequently passes with or without the fix. If reverting the fix does not turn the test red again, the test is not measuring what you think and the fix is unverified.

In practice you can collapse this into a one line instruction the agent runs itself, if your agent has terminal access and you trust it with git:

Before fixing anything:
1. Write a test reproducing the reported bug. Touch test files only.
2. Run it. Paste the failure output. If it does not fail, stop and tell me.

After I approve the reproduction:
3. Make the minimal implementation change to pass it.
4. Run the test (expect pass), then `git stash` the implementation
   change, run the test again (expect fail), then `git stash pop`.
5. Paste all three results. If step 4's second run passed, the test does
   not capture the bug: say so and go back to step 1.
6. Run the full suite.

When the bug resists reproduction

Some bugs genuinely will not reduce to a unit test on the first attempt. Handle those by widening the net rather than abandoning the rule:

Situation

What to ask for instead

Only happens in production

A test at the boundary the production data crosses, fed the real payload that broke

Depends on timing or ordering

A test that forces the order deterministically, even if it needs an injected clock

Only on one customer's data

A fixture built from that record with identifying fields scrubbed

Intermittent

A test run in a loop with a seed, plus the seed that reproduced it

The intermittent case deserves care, because it is where agents and humans alike start deleting assertions to make the red go away. If the failure is genuinely non-deterministic, how to prompt AI to fix a flaky test covers that separately, and the distinction matters: a flaky test is a broken measurement, a reproduction test is a measurement you are deliberately constructing.

If after two honest attempts no test can be made to fail, that is information. It usually means the bug is in configuration, environment, or data rather than in code, and the agent is about to spend an hour changing code that was never wrong. Stop and go looking somewhere else.

The failure modes to watch for

A test that asserts the fix, not the bug

The agent writes a test asserting the function returns the new, corrected value. That test fails before the fix for a trivial reason: the behaviour is simply different. It does not demonstrate that the original behaviour was wrong, only that it was different. Ask instead for a test asserting the user-visible property that was violated, such as the total matching the sum of the line items, rather than the specific output the fix produces.

A reproduction that quietly changes the implementation

You said test files only. The agent adds a fixture, then a small helper in the source tree, then adjusts a default so the fixture works. Now the reproduction and the fix are entangled and step 6 is meaningless. This is the general problem covered in stopping an agent from editing unrelated files, and the cheap defence is to check the diff for the reproduction commit before approving it.

Tests adjusted instead of code

The oldest one. Faced with a red test, the agent modifies the test until it is green. Requiring a revert check makes this visible, since a test bent to fit the fix will not fail on revert. It is worth reading why agents change tests to make them pass if this happens more than once, because it usually indicates the instruction is being read as make the suite green rather than fix the defect.

Why this is worth the extra step

Reproduction-first adds a few minutes per bug and removes a category of work entirely: the re-opened ticket, the fix that was never a fix, the archaeology in three weeks when the same symptom returns with a different stack trace. It also leaves an artefact. Every bug fixed this way permanently adds a test that would catch its return, which is the compounding part.

The broader habit this belongs to is refusing to accept an agent's own report of success as evidence. That is the same instinct behind checking whether your agent actually ran the tests and behind reading the diff yourself before merging. For the general tooling picture, the overview of AI coding tools covers where this fits in a wider workflow.

Frequently asked questions

How do I get an AI coding agent to reproduce a bug before fixing it?

Tell it explicitly to touch test files only in the first step, to run the new test, and to paste the failure output before proposing any implementation change. Agents follow an ordering constraint reliably when it is stated as a gate with an output you will inspect, and much less reliably when it is phrased as a preference.

What if the agent writes a test that passes before the fix?

Then the test does not capture the bug and you should reject it. This usually means the agent asserted its theory of the correct output rather than the property that was violated. Ask for an assertion on observable behaviour, such as an invariant that must hold, instead of on a specific return value.

Should I let the agent write the failing test and the fix in one go?

No, and the revert check is the reason. If both arrive together you cannot verify that the test actually detects the bug, because you never saw it fail independently. Splitting them into two reviewed steps costs a minute and is the whole value of the technique.

What do I do if the bug cannot be reproduced in a test?

Try once at a wider boundary, using the real payload or a scrubbed fixture from the affected record. If that also fails to reproduce, treat it as a signal that the problem lives in configuration, environment or data rather than code, and redirect the investigation before any code is changed.

Does this work for bugs found in production logs?

Yes, and it is where the technique pays most. Extract the exact request or payload from the log, scrub anything identifying, and make that the fixture. A reproduction built from real production input is also the most valuable regression test you will write that week.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.