AI Code Comments That Lie About the Code
An agent that edits a function and leaves the docstring untouched has produced something worse than no documentation. Here is why it happens, the four patterns to watch for, and how to make review catch them.
A wrong comment is worse than no comment, because the reader stops looking. AI coding agents produce wrong comments at a higher rate than people do, for a specific and fixable reason: they are rewarded for describing intent, they edit code in small local passes, and nothing in the toolchain checks whether the description still holds afterwards.
The fix is not to tell people to read more carefully. It is to make the drift visible.
Why agents produce comments that lie
Four mechanisms, each with a different signature in the diff.
It documented the plan, not the result. The agent wrote a docstring describing what it intended, then hit a constraint partway through and changed approach. The code moved. The docstring did not. This is the most common one and the hardest to spot, because the comment is fluent and plausible.
It edited the body and left the header. You asked for a change to one branch of a function. The agent made a surgical edit, which is exactly what you wanted, and never re-read the docstring six lines above. Human developers do this too, and the research is unambiguous that it matters: one study found that inconsistent changes are around 1.5 times more likely to lead to a bug-introducing commit than consistent ones. An agent making many small edits per hour compounds it faster.
It restated the code in English. // increment counter by one above counter += 1. Harmless when written, actively misleading after someone changes the line to counter += batchSize and leaves the comment. These comments carry no information and a permanent maintenance liability.
It invented a rationale. The worst category. The agent writes "we use a linear scan here because the collection is always small" when nothing in the code or its history establishes that. The claim is unfalsifiable at review time and load-bearing later, when someone reads it and decides not to optimise.
The underlying pattern is that comments are not compiled, not tested, and not linted, so nothing pushes back. Code and comments co-evolve in roughly 90 percent of cases, but the re-documentation often lands in later revisions rather than in the commit that caused the drift. An agent working autonomously has no later revision.
The four review rules that catch this
Reviewing every comment against every line does not scale. These four rules catch most of it in the time you already spend.
Read the docstring against the signature, not the body. Parameter lists, return types, and raised exceptions are quick to check and are where drift is most damaging to callers. If the docstring mentions a parameter that no longer exists, stop and re-read the whole function.
Treat every "because" as a claim needing evidence. Any comment asserting a reason, a performance property, or a constraint gets one question in review: how do we know that? If the answer is not in the code, the tests, or a linked issue, delete the claim or replace it with a link.
Flag unchanged comments inside changed hunks. In a diff, a comment line with no
+or-sitting inside a modified block is the highest-yield thing to look at. Most review tools show this for free once you know to look.Delete comments that restate the line below. They are pure future liability. This is a mechanical judgement, so it costs no thought.
Rules three and four are where most of the value is, and both are fast. The rest of the diff discipline is covered in reviewing an agent's git diff before merging.
Make the agent produce fewer of them
Prevention beats detection here, and it is mostly a matter of instruction placement. Put this in your AGENTS.md or equivalent project instructions rather than in individual prompts, because per-prompt rules get lost across a long session, which is the failure mode described in why an agent keeps ignoring your instructions.
Comments policy:
- Do not write comments that restate what the code does.
- Comment only non-obvious constraints, invariants, and the reason
a non-obvious approach was chosen.
- Never assert a performance or data characteristic you have not
verified in this repository. If you believe one holds, write it
as a question in the PR description, not as a comment.
- When you modify a function, re-read its docstring and update or
delete it in the same edit.The third rule is the one that pays. It converts invented rationale into a reviewable question, which is where it belongs.
Also useful: ask for the docstring update as an explicit step rather than assuming it. "Update the function and its docstring" produces a different result from "update the function", reliably enough to be worth the four extra words.
A CI check that costs an afternoon
Full comment-code consistency checking is an open research problem, so do not attempt it. A crude heuristic catches a surprising share of real cases:
For each changed function in a diff, extract its docstring.
Extract the parameter names from the current signature.
Fail if the docstring names a parameter that is not in the signature.
Warn if the signature has a parameter the docstring does not mention.
That is a tree-sitter query and about forty lines. It catches the "edited the body, left the header" case, which is the most common and most damaging. It will not catch invented rationale, and nothing automated will. That one needs the review rule.
For typed languages, much of this comes free: if your docs are generated from types, the type checker is already your consistency check. This is a real argument for generated documentation over hand-written prose, and it interacts with the approach in using AI to write documentation.
Where comments still earn their place
None of this is an argument against comments. It is an argument against a particular kind. Comments that survive contact with change tend to be the ones an agent would never think to write:
Comment type | Survives change | Why |
|---|---|---|
Restates the code | No | Breaks on the next edit |
Describes the plan | No | The plan changed |
Invented rationale | No, and it misleads | Was never true |
Links to an issue or spec | Yes | The link stays valid |
Names a real external constraint | Yes | The constraint is outside your code |
Warns about a non-obvious ordering | Usually | Encodes something the code cannot say |
The last three are what you want more of, and they are the ones you have to ask for specifically. Left alone, an agent will generate the first three, because they are what most training data contains. The rest of the workflow around this sits in the wider guide to AI coding tools.
FAQ
Should I just tell the agent to stop writing comments?
Tempting, and slightly too blunt. A zero-comment policy loses the genuinely useful ones. A better instruction is to allow only comments about constraints, invariants, and reasons, which cuts the volume sharply while keeping the load-bearing cases.
Do AI code comments actually drift faster than human ones?
The drift mechanism is identical. The rate differs because of throughput: an agent can make twenty small edits in the time a person makes two, and each edit is an opportunity for a comment to go stale. The problem is not new, only faster.
How do I fix a codebase that is already full of stale comments?
Do not do a sweeping pass, and definitely do not ask an agent to update every comment in the repository, which mostly generates confident new fiction. Fix comments in the files you are already touching, and delete rather than rewrite when you are not certain.
Is it worth asking an agent to audit its own comments?
Somewhat, if you scope it. "List every comment in this diff that makes a factual claim about behaviour, and for each say what in the code supports it" produces a reviewable list. Asking "are these comments correct?" produces reassurance.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


