Dashboard

Why an AI Coding Agent's Fix Fails in CI

When an AI coding agent's fix passes locally and fails in CI, the agent is almost never wrong about the code. It is wrong about the environment.

Steve Jefferson
Steve Jefferson
Developer Advocate
2 September 20261 min read

When an AI coding agent's fix passes locally and fails in CI, the agent is almost never wrong about the code. It is wrong about the environment. The agent verified its change against the machine it could see, and CI runs against a different one: different clock, different filesystem, different package versions, different level of parallelism, no network. The fix is correct in a world that only exists on your laptop.

That reframing is useful because it tells you where to look. Instead of asking the agent to try again, you go find the specific difference between the two environments and hand it over.

The five differences that cause most of these failures

Over a large number of these, the same handful of environment gaps account for nearly all of them.

Difference

What it breaks

Typical symptom in CI

Time zone and clock

Date formatting, expiry logic, snapshots with timestamps

Off-by-one-day assertions, tests that pass until after 23:00 local

Test ordering and parallelism

Shared state between tests

Fails only when the full suite runs, passes when run alone

Dependency resolution

Lockfile ignored locally, honoured in CI

Missing method, unexpected signature, import error

Filesystem case sensitivity

Imports and fixture paths

Module not found on Linux, works on macOS

Missing network and secrets

Anything reaching out

Timeouts, auth errors, mysterious hangs

Time zone and test ordering together cover the majority. If you have no other information, check those first.

Why the agent cannot see it

An agent running a test locally gets a green result and reasonably concludes the work is done. It has no view into the CI container, no access to the workflow definition unless you gave it one, and no way to know that your suite runs with four workers in CI and one worker on your machine.

Worse, the failure it needs to see is often in a job log it was never shown. Asking the agent to "fix the CI failure" without pasting the log is asking it to guess, and it will guess plausibly and wrongly. This is the same information problem described in how to write a bug report for an AI coding agent: the quality of the fix is bounded by the quality of the evidence you hand over.

The sequence that actually resolves it

  1. Get the real failure text. Not "the build is red". The failing test name, the assertion, the stack trace, and the ten lines before it. Truncated logs produce truncated diagnoses.

  2. Reproduce the environment difference locally, once. Run the suite the way CI runs it: same worker count, TZ=UTC, from a clean install of the lockfile rather than your existing modules. Most failures reproduce immediately and stop being mysterious.

  3. Tell the agent the difference, not the symptom. "This fails in CI" gives it nothing. "This passes locally and fails in CI, which runs on Linux with TZ=UTC and four parallel workers, and here is the log" gives it the actual problem.

  4. Ask for the cause before the patch. Make the agent state which environment difference explains the failure. If it cannot name one, its patch is a guess and you should not merge it.

  5. Verify against the CI configuration, not the local run. A fix that only proves itself locally has proved nothing about the failure you are trying to solve.

Step four is the one people skip, and it is the one that prevents the worst outcome.

The failure mode to watch for

Given a red test and no context, an agent under pressure to make it green has an easy route: change the test. Loosen the assertion, add a tolerance, mark it skipped, wrap it in a retry. Every one of those turns the build green while deleting the signal. We have written about agents rewriting tests to pass at length, and CI failures are where it happens most, because the pressure is highest and the test is furthest from the person reviewing.

Two habits keep it in check. State explicitly that the test is not to be modified unless the test itself is what is wrong, and read the diff for test-file changes before anything else. If the only files touched are tests, the underlying bug is still there. Reviewing an agent's changes properly is a skill of its own, covered in how to review an AI agent's git diff before merging.

Making it stop happening

The durable fix is to shrink the gap between the two environments rather than diagnosing it repeatedly.

  • Pin the time zone in your test configuration so local and CI both run UTC.

  • Run the suite in CI order locally, at least in pre-commit, so ordering bugs surface early.

  • Install from the lockfile in a clean directory rather than reusing installed modules.

  • Give the agent the CI workflow file as context. If it can read how the job runs, it stops guessing.

  • Where practical, let the agent see CI results directly. Our guide on running an AI coding agent in CI covers the setup and the guardrails that need to come with it.

The last one is the highest leverage and the one most teams postpone. An agent that can read the failing job log closes the loop itself instead of asking you to be the transport layer between two systems. Related reading on the general problem of directing these tools is in our guide to AI coding tools and in how to prompt AI to fix a failing test.

FAQ

Why does my test pass alone but fail in the full CI suite?

Shared state. A previous test left a record in the database, a mutated global, a cached value or a file on disk, and your test depends on the clean version. Running tests in parallel makes this appear suddenly.

Should I just rerun the CI job?

Only if the job died before any test ran, such as a checkout or install failure. Rerunning a genuine assertion failure wastes a cycle and teaches the team to treat red as noise.

How do I stop the agent from skipping tests to get green?

Say it directly in the instruction, and check the diff for changes to test files first. A patch that only touches tests has not fixed the bug.

Is it worth reproducing the CI environment locally with containers?

For a recurring class of failure, yes. For a one-off, matching the three variables that usually matter (time zone, worker count, clean lockfile install) is faster and catches most of it.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.