Dashboard

Why AI Coding Agents Leave TODO Comments Everywhere

The mechanistic reason AI coding agents park uncertainty in TODO comments instead of resolving it, and the before/after prompt pattern that stops it.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
20 September 20261 min read

AI coding agents leave TODO comments because they're built to produce output that looks finished within one task turn, and a TODO is the cheapest way to mark a decision the model isn't confident enough to make right now without stopping entirely. It isn't laziness. It's a mix of three things: the model treating uncertainty as something to park rather than resolve, the economics of a limited context window that reward finishing the visible task over chasing every edge case, and training that rewards output which reads as complete more than output that is actually complete. Once you understand which of the three is driving a given TODO, whichever AI coding tool you use, the fix is a prompt pattern, not a plea.

Parking Uncertainty Instead of Blocking On It

When a model hits a genuinely ambiguous decision mid-task, such as which error-handling strategy an unwritten API endpoint should use, it has two options: stop and ask, or make a call and flag it. Most agentic setups are tuned to keep moving rather than interrupt the user constantly, since a coding agent that asks a clarifying question every few minutes is worse to use than one that makes reasonable assumptions and keeps going. A TODO comment is what "keep going but flag it" looks like in code. It's a deferred decision, not a skipped one, and the model is often quite capable of resolving it if you go back and ask directly instead of leaving it parked. This is the same underlying tension that shows up when an agent runs out of context mid-task: pushed past what it can reliably track, it starts parking decisions rather than making them well.

Context Window and Task-Scoping Economics

Every agent operates inside a finite context window, and every additional file it reads, every extra branch of logic it fully implements, costs budget it could spend elsewhere in the same task. Anthropic's own documentation on how context windows work describes the window as holding the full conversation, every file read, and every command output, which makes it a genuinely scarce resource within a session, not just a soft limit. Fully implementing a rarely-hit edge case, writing its tests, and wiring it into three other files is expensive in that budget. Writing // TODO: handle the empty-cart case is nearly free. When a task has twelve things that could be done and budget for eight, a model that's been trained to be helpful within its limits will often do the eight it's confident about and mark the other four, rather than doing four completely and silently dropping the rest with no trace at all. The TODO is, in a narrow sense, the model being honest about its own scoping decision.

Training Incentive Toward "Looks Complete" Output

Models are shaped heavily by human and automated feedback on whether an output looks like a reasonable, finished response. A function with a TODO comment and a working happy path reads, at a glance, as further along than a function that throws an explicit "not implemented" error or one that's missing entirely. That asymmetry gets reinforced across enough training signal that the safer move, in the model's learned sense of "safer," is to always produce something that resembles a complete implementation rather than something that visibly stops. A TODO comment is a low-cost way to keep the surface area looking finished while quietly admitting a gap exists to anyone who reads carefully. It's a compromise between two things the model is optimized for at once: appearing complete, and not silently fabricating behavior it can't actually implement correctly.

The Before/After Prompt Pattern That Eliminates It

Before: "Implement the checkout flow." Handed a broad, open-ended task like this, an agent has to make dozens of small scoping calls with no guidance on which ones you'd actually want it to stop and ask about, so it defaults to marking the ambiguous ones and moving on. The task itself never told it that TODOs were unacceptable, or which specific parts mattered enough to deserve a real decision instead of a placeholder.

After: "Implement the checkout flow. Do not leave TODO, FIXME, or placeholder comments. If you hit a decision you're not confident about, such as how to handle a declined payment or an empty cart at submit time, stop and list those decisions to me before writing code for that part, instead of guessing and flagging it in a comment. Everything you do write must be complete and tested, not partially implemented." This works because it removes the model's cheapest escape hatch and replaces it with an explicit, higher-value one: surfacing the question to you directly, before code exists, rather than after. It also sets a concrete bar, tested and complete, which is much harder to satisfy with a placeholder than a vague instruction to "finish the feature" is.

The same pattern helps when you're taking over a task an agent left unfinished: ask it to list every TODO and placeholder it left, with the specific decision behind each one, before you start reviewing the diff. That turns an unreadable scatter of comments into a short list of actual open questions, which is a much faster review than grepping the codebase for TODO after the fact.

Do all AI coding agents leave TODO comments the same amount?

No. It varies by model and by how the task was scoped. Narrower, well-specified tasks produce far fewer TODOs than broad, open-ended ones, because there's less ambiguity for the model to park in the first place. It's the same reason a model that keeps reintroducing a bug it already fixed often does so on the same kind of vague, broadly scoped tasks that produce the most TODOs.

Is a TODO comment ever the right output from an AI coding agent?

Sometimes, for genuinely out-of-scope work you didn't ask it to touch, such as a note that a separate legacy module could also use the same refactor. The problem is TODOs standing in for decisions inside the task you did ask for.

How do I stop an AI coding agent from leaving TODOs in code it already wrote?

Ask it to grep the diff for TODO, FIXME, and similar markers, list the decision behind each one, and resolve them one at a time with you before you merge, rather than asking it to "clean up" in one vague pass.

Does telling an AI agent 'no shortcuts' actually work?

Rarely on its own, since it gives the model no concrete bar to meet. Pairing it with an explicit instruction to stop and ask about specific ambiguous decisions, instead of guessing and flagging them, works far better because it replaces the escape hatch with a cheaper alternative the model can actually take.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.