Dashboard

Letting an AI Coding Agent Touch Your Auth Code

Yes, you can let an AI coding agent touch your auth code, and the honest answer is that most teams already do. The question worth asking is not whether to allow it but what has to be true before the change merges.

Steve Jefferson
Steve Jefferson
Developer Advocate
2 September 20261 min read

Yes, you can let an AI coding agent touch your auth code, and the honest answer is that most teams already do. The question worth asking is not whether to allow it but what has to be true before the change merges. Auth is not special because agents are bad at it. It is special because the failure is silent: broken authentication does not throw an error, it returns 200 and lets the wrong person in.

Every other bug announces itself. This one does not, which is why the review process matters more here than anywhere else in the codebase.

What agents get right and wrong in auth code

Worth being specific, because the risk is not uniform.

Agents are reliably good at the mechanical parts: wiring up a library correctly, adding a middleware in the right place, writing the session cookie flags, following an OAuth flow that is well documented. This is pattern-matching against a large body of public examples and it works.

Agents are unreliable at the parts where correctness depends on your application's specific model of who may do what:

  • Authorisation, as opposed to authentication. Confirming who someone is has one right answer. Deciding what that person may see depends entirely on your data model, and an agent inferring it from surrounding code will guess.

  • The negative cases. A change that handles the logged-in path correctly and quietly widens the logged-out path passes every test you wrote, because you wrote tests for the path you were thinking about.

  • Multi-tenant boundaries. The single most common serious defect is a query that filters by resource id and forgets to also filter by tenant or owner. It works perfectly in every test where one account exists.

  • Token lifetime and revocation. Refresh logic, expiry handling and "log out everywhere" are easy to write plausibly and hard to write correctly.

The pattern is consistent: agents do well where the answer is in the library documentation, and poorly where the answer is in your head.

Scope the change before you start

The best control is applied before the agent runs, not after.

Keep auth changes in their own commit, touching only auth files. Mixing an auth change into a larger feature branch is how a permission regression rides in unnoticed under forty files of unrelated diff. Our note on which files an AI coding agent should never write covers where to draw those lines generally, and auth belongs in the narrowest bracket.

State the security model in the prompt rather than hoping it will be inferred. Two sentences is usually enough:

Every query in this file must filter by both resource id and the
current user's organisation id. A user may only read records
belonging to their own organisation. Do not add a code path that
returns data when the session is absent or invalid.

That is a specification. Without it, the agent optimises for the request you made and treats the security property as an implementation detail it may reshape.

Finally, do not hand over the secrets. Agents do not need real keys to write auth code, and a real credential in a prompt is a credential in a log somewhere. See whether an AI coding agent can leak your API keys for how that goes wrong in practice.

The review checklist that catches the real bugs

Read the diff with these six questions. They are ordered by how often each one finds something.

  1. Does every data query filter by owner as well as by id? Search the diff for findById, where id = and equivalents. Each one needs a second condition.

  2. What happens with no session? Trace the path where the session is missing, expired or malformed. Confirm it denies rather than falling through to a default.

  3. Did any check move earlier or later? A permission check relocated below the point where data is fetched still returns the data.

  4. Are the cookie and token flags unchanged? httpOnly, secure, sameSite and expiry get rewritten by well-meaning refactors.

  5. Did an error message get more helpful? "No user with that email" tells an attacker which addresses are registered. Generic failures are correct here.

  6. Is anything now trusting client input? A role, tenant id or permission flag read from the request body rather than the session is the classic privilege escalation.

Run these against the diff itself rather than against the agent's summary of the diff. The summary describes intent; the diff describes behaviour. Reviewing an agent's git diff properly is the general version of this discipline.

Test the cases you did not ask for

The tests an agent writes cover the scenario in the prompt. The bugs live outside it. Add these by hand, once, and keep them:

  • A user from organisation A requesting a resource belonging to organisation B, expecting a 404 or 403.

  • A request with no credentials at all against every protected route.

  • A request with a valid but expired token.

  • A request with a valid token for a deleted or deactivated user.

The second one is worth automating across your whole route table. A single test that enumerates every route and asserts an unauthenticated request is rejected will catch the next accidental exposure without anyone remembering to look for it.

For the broader class of security defects agents introduce, see how to catch an AI coding agent introducing a vulnerability, and for the general practice of directing these tools, our guide to AI coding tools.

FAQ

Is it safe to let an AI coding agent write authentication code?

For standard, well-documented flows, generally yes, provided the change is isolated, the security model is stated in the prompt, and the diff is reviewed against the checklist above. Authorisation logic needs more scrutiny than authentication logic.

What is the single most common auth bug agents introduce?

A query filtered by resource id without also filtering by owner or tenant. It passes every test in which only one account exists, and it exposes other customers' data in production.

Should the agent have access to my auth secrets?

No. It can write the code without them. Use placeholders and inject real values from your secret store at runtime.

Do I need to review auth changes differently from other changes?

Yes, because auth failures do not raise errors. A broken feature is visible in a minute; broken authorisation returns a normal response and can run for months. Related: adding two-factor authentication to an AI-built app.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.