Dashboard

What to Do When an AI Coding Agent Says It's Done

Completion is a claim from an interested party. Seven checks in yield order, plus the three buckets that decide how you respond to a failure.

Steve Jefferson
Steve Jefferson
Developer Advocate
25 September 20261 min read

An agent reporting completion is reporting its belief about its work, produced by the same process that produced the work. It is not an independent check, and it fails in a correlated way: the misunderstanding that caused a bad implementation also causes a confident report that the implementation is good. So when an AI coding agent says it's done, the useful question is not whether to verify but what to check first, because your attention is the scarce resource and the checks differ enormously in yield.

Here is the order that finds the most problems the fastest. Run them in sequence and stop at the first real failure, because one failure usually predicts others.

1. Read the diff, not the summary

Thirty seconds, highest yield of anything on this list. The summary is generated prose about the work. The diff is the work. They disagree more often than you would expect, and almost always in the same direction: the summary describes the intended change, the diff contains the intended change plus something else.

Specifically scan for files you did not expect to be touched. An agent fixing a test that reaches into configuration, a feature change that also reformats an unrelated module, a dependency version quietly bumped. None of these appear in a summary because the agent does not consider them part of the task it is reporting on.

2. Check that the tests ran, not that they exist

"I added tests and they pass" is the single least reliable completion claim, and it fails in a way that is invisible unless you look for it. Tests can be added without being run, run without being included, or reported as passing from a partial run.

Run them yourself. It costs one command and it is the difference between believing a claim and having a result. If you want the longer version of this specific failure, checking whether the agent actually ran the tests covers how to tell after the fact.

3. Look for the weakened assertion

This is the one that costs real money later. When an agent cannot make code satisfy a test, changing the test is available to it and often easier than fixing the code. The result is green CI and a test that no longer tests anything.

The specific patterns, worth knowing by sight:

  • An exact equality check replaced with a truthiness or not-null check.

  • An assertion on a value replaced with an assertion on a type.

  • A test case deleted rather than fixed, with no mention in the summary.

  • A skip, xfail, or conditional-skip marker added to an existing test.

  • An error-handling branch turned into a broad catch that swallows the failure.

Grep the diff for the skip markers your framework uses before anything else. Most frameworks document these plainly, for example pytest's skip and xfail markers, and a grep takes two seconds. It catches the worst version of this immediately.

4. Verify the thing you actually asked for

Agents are strong at the tractable adjacent problem and will sometimes solve that instead. You ask for pagination on an endpoint and get a tidy, well-tested pagination helper that nothing calls. Everything about it is correct except that your endpoint is unchanged.

The check is to go back to your original request, take the specific observable behaviour you wanted, and confirm that behaviour exists. Not that code supporting it exists. Call the endpoint. Click the button. Run the command. This is where a lot of otherwise careful reviews stop one step short.

5. Check the edges the agent had no reason to consider

Agents optimise for the described case. Absent instruction, the unspecified cases get whatever behaviour falls out.

Case

Typical agent handling

Check

Empty input

Unhandled or crashes

Pass an empty list or string

Concurrent calls

Assumed serial

Fire two at once

Partial failure midway

No rollback

Kill it halfway through

Very large input

Loads everything

Try 10,000 rows

Repeat of the same call

Not idempotent

Run it twice

Run it twice is the highest-value single check here and takes seconds. A surprising share of agent-written handlers work perfectly once and do something unpleasant the second time, because nothing in the task description mentioned that it might be called more than once.

6. Look at what was deleted

Additions get reviewed. Deletions slide past, because a diff with a lot of green draws the eye to the green. Filter the diff to removals only and read them on their own.

You are looking for a guard clause that is gone, a validation step that was in the way, a comment that documented a non-obvious constraint, or a fallback branch that seemed redundant. Agents remove things that look unnecessary from inside the immediate task, and the whole point of a guard clause is that it looks unnecessary until the day it is not.

7. Ask it what it is unsure about

Cheap, and it works better than asking it to check its own work. A request to verify tends to produce a confirmation, because the model is generating from the same understanding that produced the code. A request to name uncertainties and assumptions produces something more useful, because it does not require contradicting itself.

Phrase it as "what assumptions did you make that I did not specify, and which part of this are you least confident in?" The answers point at the edges of what it understood, and those edges are where the defects are. If it keeps making the same class of mistake across tasks, that is a different problem, and an agent that reintroduces the same bug needs a fix at the instruction level rather than another review pass.

What to check first when an AI coding agent says it's done

Not all completion claims are equally risky. Ranked by how often they turn out to be wrong:

  1. "Tests pass." Verify by running them. Most commonly wrong, cheapest to check.

  2. "I did not change anything else." Verify with the file list in the diff.

  3. "This handles the edge cases." Usually means the ones that were named.

  4. "This is backwards compatible." Rarely actually exercised against an old caller.

  5. "I refactored it for clarity." The claim most likely to hide unrelated behaviour change.

None of this means agents are unreliable in a way that makes them not worth using. It means completion is a claim from an interested party, and the review has to be structured rather than vibes-based. Fifteen minutes in this order beats an hour of unstructured reading, and it beats trusting the summary by a margin that shows up in production.

What to do when a check fails

Finding a problem is the easy part. The instinct is to point it out and ask for a fix, which works for a narrow defect and goes badly for anything structural, because the agent patches the specific thing you named while the misunderstanding that produced it stays in place. You then get a second completion claim built on the same faulty model of the task.

Sort the failure into one of three buckets first, because the right response differs:

  1. A local defect. One wrong line, one missing case, nothing conceptually off. Point at it and let the agent fix it. This is the cheap case and most failures are here.

  2. A misunderstanding of the goal. It built the wrong thing competently. Do not iterate on the wrong thing. Revert, rewrite the task description with the ambiguity removed, and start again, because patching toward a different target costs more than restarting.

  3. A trust failure. Weakened assertions, deleted tests, unrelated files touched. Revert the whole change without negotiating. The issue is not this diff, it is that the reported state and the real state came apart, and anything else in that diff is now unverified.

The second bucket is the one people handle worst. Sunk cost makes a mostly-working wrong implementation feel closer to done than an empty branch, and it usually is not. If the agent misread the goal, every subsequent instruction is a correction layered on a bad foundation, and you can spend an hour steering toward something a clean restart would have reached in ten minutes.

Whatever the bucket, feed the fix back into the instructions rather than only into the conversation. A constraint that exists only in a chat turn is gone next session, which is how teams end up making the same correction weekly. Constraints that matter belong in the repo, in the agent instructions file, or in a test that fails loudly, and knowing which files the agent should never write is a good place to start writing them down.

FAQ

Why do AI coding agents say they are done when they are not?

Because the completion report is generated by the same process as the work, so any misunderstanding is reflected in both. The agent is accurately reporting its belief. That belief is not independent evidence, which is what makes self-reported completion weak.

What is the single fastest check?

Read the file list in the diff. Unexpected files are the strongest early signal that the agent did something beyond the task, and it takes a few seconds. Grepping for newly added test skip markers is a close second.

Should I just ask the agent to double check its work?

It helps less than people expect, because reviewing its own output means reasoning from the same understanding that produced it. Asking what assumptions it made and where it is least confident gets more useful answers than asking whether the work is correct.

How much of this can be automated?

Most of steps two, three, and five. Running the test suite yourself, failing the build on newly added skip markers, and having a few adversarial edge case tests that always run will catch a large share of false completions without any human attention. The broader tooling landscape is covered in our guide to AI coding tools.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.