Dashboard

Take Over a Task Your AI Coding Agent Didn't Finish

When an agent stops halfway, do not read its summary and do not start over. Read the diff and build a three-column audit of what changed, what was verified, and what was assumed.

Steve Jefferson
Steve Jefferson
Developer Advocate
16 September 20261 min read

Take Over a Task Your AI Coding Agent Didn't Finish

When an AI coding agent stops halfway, do not read its summary and do not start over. Read the diff and build a three-column audit of what changed, what was verified, and what was assumed. That audit takes about ten minutes on a typical half-finished task and it is the only thing that tells you whether you are 70% done or 20% done wearing a 70% costume.

Agents stop mid-task constantly: a usage limit, a crashed session, a tool that timed out, or you hitting escape because it was heading somewhere wrong. The code it already wrote is still on disk. The question is what to trust.

Why the agent's own summary is the wrong starting point

An agent's closing message describes what it intended to do. That is not the same as what it did, and the divergence is worst exactly when the run was cut short, because the summary is often generated from the plan rather than from the result.

The specific traps:

  • It reports a file as "updated" when the edit failed and it moved on.

  • It reports tests as passing when it ran a subset, or ran them before the last three edits.

  • It describes a function it planned to write in a step it never reached.

  • It omits files it touched incidentally, which is where most surprises live.

None of this is the agent lying. A summary is a forecast written in the past tense. Treat it as a hypothesis and verify it against the working tree.

The ten-minute handoff audit

Do these in order. The whole point is to build the artifact neither side produced.

1. Get the real change set

bash
git status --short
git diff --stat
git diff

--stat first gives you the shape: how many files, how much churn, and whether anything unexpected is in there. If the agent touched a file that has nothing to do with the task, that is your first flag and it is worth reading closely. Agents drift into adjacent files more often than their summaries admit.

If the agent worked on a branch, compare against the base rather than the last commit, or you will miss everything it committed along the way:

bash
git diff main...HEAD --stat

2. Sort every changed file into three columns

This is the artifact. A scratch file is fine.

Changed

Verified

Assumed

The edit exists in the diff

Something proved it works

The agent relied on it being true

For each file, ask what evidence exists that the change is correct. A test that exercises it counts. A type checker passing counts for the narrow thing types check. The agent saying "this should work" does not count and goes in the third column.

The third column is the valuable one, and it is what neither a summary nor a diff gives you on its own. Typical entries: an API returns the shape the agent expected, a column exists in the database, a config value is set in the environment it did not have access to, an upstream function is safe to call concurrently.

3. Run the tests yourself, all of them

Not the ones the agent ran. Yours, from a clean state.

bash
git stash list
npm test 2>&1 | tail -40

Check the stash list first. Agents sometimes stash work and lose track of it, which is a silent way to have a change that exists but is not in your diff. The git stash documentation covers recovering entries that were dropped rather than popped.

If the suite fails, resist fixing it immediately. Note which failures are from the agent's changes and which were already broken. Running the suite on the base commit takes thirty seconds and tells you which is which:

bash
git stash && git checkout main && npm test; git checkout - && git stash pop

4. Find the actual stopping point

The last edited file is usually not where the work stopped conceptually. Look for the seam: a function called but never defined, an import added for something not yet written, a TODO the agent left, a test file created with one test in it.

bash
git diff | grep -n "TODO\|FIXME\|NotImplemented\|pass$"

That seam is where you pick up. Everything before it is code you now own and need to have read.

5. Decide: continue, or restart from a better prompt

Now you have enough to choose, and the choice is usually clear from the audit rather than from feeling.

Continue when the changed set is coherent, the assumed column is short, and the seam is obvious. Restart when the diff sprawls across unrelated files, or when the assumed column has three or more entries you cannot cheaply check. A sprawling half-change is more expensive to reason about than an empty file, and reverting is not an admission of failure. It is a cost decision.

If you restart, do not throw the audit away. It is the best prompt you will write today, because it contains exactly the constraints the first attempt lacked.

Handing back to the agent instead of finishing yourself

Sometimes the right move is to resume the agent with better information. The handoff works in that direction too, and the audit is what you feed it.

A resume prompt that works:

text
Partial work exists in the working tree. Do not start over.

Already done and verified by me:
- src/billing/invoice.ts: new lineItems() helper, covered by invoice.test.ts

Already done, NOT verified:
- src/billing/tax.ts: rate lookup added, no test exists

Assumptions I have not checked:
- getTaxRate() is safe to call per-line-item rather than per-invoice

Stopped at: calculateTotal() calls applyTax(), which does not exist yet.

Next: write applyTax() only. Do not modify invoice.ts or lineItems().

The three things that make this work: an explicit list of what is off limits, the assumptions surfaced as assumptions, and a single next step. Without the "do not modify" line, an agent picking up a partial change will frequently rewrite the part you already verified, and you will audit it twice.

This is the same discipline as reviewing an agent's git diff before merging, applied earlier in the cycle. It is also why keeping an agent from losing context matters: the less context evaporates when a run ends, the smaller the audit.

Make the next interruption cheaper

A few habits that turn a ten-minute audit into a two-minute one:

  • Commit at the seams. Ask the agent to commit after each self-contained step. A half-finished task with four commits is dramatically easier to reason about than one with none. Using git with AI coding agents covers the workflow.

  • Work on a branch, always. It makes the base-versus-head comparison trivial and the revert free.

  • Require a plan before edits. If the plan is on screen, you know what the agent was about to do next without inferring it from a dangling import.

  • Keep tasks small enough to finish in one run. The real fix for interrupted work is work that fits.

For the broader workflow these fit into, see how AI coding tools fit together. And if the audit turns up changes you want gone rather than finished, rolling back a bad agent change is the cleaner path.

One last note on verification: if the agent claims it ran tests before it stopped, that claim deserves the same scepticism as the rest of the summary. Checking whether it actually ran them is a thirty-second check that regularly changes the answer.

FAQ

Should I just revert and start over?

Sometimes, and the audit tells you. Revert when the diff spans unrelated files or when you cannot cheaply verify the agent's assumptions. Continue when the change set is coherent and the stopping point is obvious.

How do I know which files the agent actually changed?

git status --short and git diff --stat against the base branch. Do not rely on the agent's summary, which reflects intent rather than outcome, and check git stash list for work that is not in the diff.

Can I just tell the agent to continue?

Yes, but give it the audit rather than "keep going". Name what is verified, what is not, what you are assuming, and which files it must not touch. Without that last constraint it will often rewrite work you already checked.

What if the tests were already failing before the agent ran?

Run the suite on the base commit to separate pre-existing failures from new ones. Thirty seconds of checking prevents an hour of debugging something the agent never touched.

Why not trust the agent's summary at all?

Trust it as a hypothesis about what was attempted. A summary is written from the plan, and an interrupted run is exactly the case where the plan and the result diverge most.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.