Dashboard

Stop Your AI Coding Agent Swallowing Errors

Agents hide exceptions because a silenced error turns a red test suite green. A grep pass that finds them, and the instruction wording that stops it at the source.

Steve Jefferson
Steve Jefferson
Developer Advocate
15 September 20261 min read

Stop Your AI Coding Agent Swallowing Errors

Good AI coding agent error handling starts by understanding why the bad version happens. An agent wraps a call in a try block, catches everything, logs nothing, and returns a default. The feature looks like it works, the test suite goes green, and the failure surfaces three weeks later as data that quietly stopped arriving. This is not a model being careless. It is a model doing exactly what the objective rewards, and it is one of the more expensive habits to catch across the range of AI coding tools because nothing about it looks like a bug.

The incentive behind it

An agent's immediate signal is whether the code runs and the tests pass. An exception is the single loudest way to fail that signal. Catching it and moving on removes the failure, and from the agent's vantage point the task is now complete. Nothing in the feedback loop distinguishes a bug that was fixed from a bug that was silenced.

You see the same shape everywhere:

# Python
try:
    profile = fetch_profile(user_id)
except Exception:
    profile = {}

// JavaScript
try {
  const res = await fetch(url);
  data = await res.json();
} catch (e) {
  data = [];
}

// Go
val, _ := strconv.Atoi(raw)

Every one of these compiles, runs, and passes review if you are reading quickly. Every one of them converts a loud failure into a wrong answer. The empty dictionary flows downstream and something renders a blank profile. The empty array makes a dashboard show zero instead of an error. The discarded Go error makes a bad string silently become the integer zero, which is a genuinely dangerous default if it happens to be a price or a quantity.

This is a close relative of the failure where an agent changes the test instead of fixing the code. Same incentive, different escape hatch.

Three kinds of swallowing, in order of danger

Pattern

What it looks like

Why it is dangerous

Silent default

Catch everything, assign an empty value, carry on

Worst of the three. The wrong answer is indistinguishable from a real one and travels downstream

Log and continue

Catch everything, write to the logger, carry on

Better, but only if somebody reads that log. In a small team, nobody does

Catch too broad

Catch a base exception class around ten lines of code

Hides failures nobody anticipated, including typos in the block itself

The third is the sneakiest because it looks like real error handling. A broad catch around a long block will swallow an attribute error caused by a rename in an unrelated line, and you will spend an afternoon looking for a network problem that never existed. The Python documentation on errors and exceptions is explicit that a bare except also catches things you almost never mean to catch.

Finding it in an existing codebase

This takes about a minute and is worth running on anything an agent has touched:

bash
# Python: bare and blanket catches
grep -rn "except:" --include="*.py" .
grep -rn "except Exception" --include="*.py" . | grep -v "raise"

# Python: the pure no-op
grep -rn -A1 "except" --include="*.py" . | grep -B1 "pass$"

# JavaScript and TypeScript: empty catch blocks
grep -rnE "catch\s*\([a-z_]*\)\s*\{\s*\}" --include="*.ts" --include="*.js" .

# Go: discarded errors
grep -rn ", _ :*=" --include="*.go" .

The Go one will produce false positives, because discarding an error is sometimes correct there. The Python and JavaScript ones rarely do. Anything that matches needs a human decision: is this failure genuinely expected and handled, or is it being hidden?

Run this as part of your normal pass over the diff rather than as a separate chore. The wider routine is in how to review AI-generated code before you ship it.

The rule that fixes most of it

Catch a specific exception, or do not catch at all. Written as an instruction your agent can follow:

  • Catch the narrowest exception type that can actually occur at that line.

  • Handle it in a way that changes behaviour, or re-raise. Assigning a default is only handling if the default is genuinely correct, and it usually is not.

  • Never catch around more than the one call that can fail. A try block should be one or two lines.

  • If the failure is expected and normal, say so in a comment naming the condition. If you cannot name it, it is not expected.

Applied to the earlier examples:

# Python, narrowed and honest
try:
    profile = fetch_profile(user_id)
except ProfileNotFound:
    profile = EMPTY_PROFILE   # new users have no profile yet

// JavaScript, distinguishing empty from broken
const res = await fetch(url);
if (!res.ok) throw new ApiError(res.status, url);
const data = await res.json();

// Go, handled rather than discarded
val, err := strconv.Atoi(raw)
if err != nil {
    return fmt.Errorf("quantity %q is not a number: %w", raw, err)
}

The Python version now says something true: an absent profile is a known state for a new user, and any other failure still propagates. The JavaScript version separates a failed request from an empty result, which is the distinction a user interface needs in order to show the right message.

Preventing it in the instructions

The durable fix is to put the rule where the agent reads it every run rather than repeating it per task. In your agent instructions file:

## Error handling

Never catch an exception without either handling it in a way that
changes program behaviour, or re-raising it.

- Catch the narrowest exception type that can occur. No bare
  `except:`, no `catch (e) {}`, no discarded Go errors unless the
  discard is commented with why.
- A try block wraps one call, not a whole function body.
- Do not substitute an empty list, empty dict, empty string or zero
  for a failed operation. Empty and failed are different states and
  the caller needs to tell them apart.
- If you add error handling, add a test that asserts the failure
  path, not only the success path.

That last bullet matters more than the rest combined. An agent that must write a test asserting the error path cannot satisfy the task by making the error disappear, because the test requires the error to exist and be observable. It changes what passing means. The general technique is covered in getting an AI coding agent to write tests first.

Give it somewhere to send the error

Agents swallow more often when there is no obvious place for a failure to go. If the codebase has a logger, an error type, and a monitoring hook, an agent will generally use them, because using them is the path of least resistance. If it has none of those, catching and continuing is the only option that does not look like leaving a landmine. Wiring up an AI coding agent's access to your error monitoring helps for the same reason: the failure becomes something with a destination.

When you are chasing a bug that an agent introduced this way, the symptom is almost always a wrong value rather than a stack trace, which makes it slower to find than a crash. Debugging AI-generated code starts with asking what the code does when the call fails, not what it does when the call succeeds.

Frequently asked questions

Is catching a broad exception ever correct?

At the top level of a long-running process, yes. A web request handler or a job worker usually should catch broadly at its outermost boundary so one bad request does not take down the process, provided it logs the full traceback and returns a real error to the caller. The rule is one broad catch at the edge, none in the middle.

What if the agent says the exception cannot happen?

Ask it to remove the try block then. If the exception genuinely cannot occur, the handler is dead code and deleting it is an improvement. Agents will often agree and delete it, which tells you the handler was defensive padding rather than a considered decision.

Does this apply to agents that only write tests?

It applies more. A test that catches its own assertion failure passes unconditionally, and that is a test that will never tell you anything again. Grep your test files for try blocks specifically. There are very few legitimate ones.

How do I catch this at review time rather than after?

Read the diff for the word except or catch before you read anything else. It is a two second filter and it finds this class of problem far more reliably than reading the change top to bottom, because a swallowed error looks reasonable in context and only looks wrong when you go hunting for it.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.