Dashboard

Which Parts of Your Codebase to Let an AI Agent Touch First

A staged rollout order for giving an AI coding agent access to your codebase, starting with what is safe to get wrong and ending with what is not.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
19 September 20261 min read

Which Parts of Your Codebase to Let an AI Agent Touch First

Give a new AI coding agent access in the same order you would give a new hire access: start where a mistake is cheap and visible, and only extend reach once it has earned it. In practice that means tests and internal tooling first, then well-covered feature code, then anything touching money, auth, or data migrations, and only once you have watched it work on the earlier tiers.

The rollout order, and why it is in this order

Tier

Example areas

Why this tier, why this order

1. Safe to get wrong

Test files, internal scripts, documentation, non-customer-facing tooling

Mistakes here are caught by CI or a human before they ship anywhere, and you learn the agent's habits at zero cost

2. Well-covered feature code

UI components, API endpoints with existing test coverage

Coverage acts as a safety net, so you are watching for review-worthy issues rather than catching untested regressions

3. Under-covered feature code

Older modules, code nobody has touched in a while

Worth pairing with a request to write tests first, since this is where an agent's confidence and your test coverage are both weakest

4. Sensitive by default

Auth, payments, permissions, data migrations

Errors here are expensive, sometimes irreversible, and often invisible until a customer hits them

What “earning it” actually looks like

There is no fixed number of pull requests that unlocks the next tier. What you are actually watching for is whether the agent's mistakes in the current tier are the kind you expected (small, caught by tests, easy to explain) rather than the kind that surprise you (silent scope creep, tests it wrote to pass rather than to verify, confidently wrong reasoning about why something works). If mistakes are getting more surprising rather than less as you extend its reach, that is a reason to hold at the current tier, not push forward on schedule.

This is different from asking how much it should see

A related but separate question is how much of the codebase an agent needs in context to work well, which is mostly a technical constraint about window size and retrieval quality. Our piece on how much of your codebase an AI coding agent should see covers that. This piece is about a policy decision: given that it can see and edit anything you point it at, what should you actually authorize it to touch, and in what order, regardless of how much context it technically has access to.

The sensitive tier deserves its own rule, not a graduation

Auth code specifically is worth treating as a standing exception rather than something that becomes fair game once trust is established elsewhere. Our dedicated piece on letting an AI coding agent touch your auth code goes into why this tier resists the same earned-trust logic that works for everything above it: the failure mode is not “more bugs,” it is “silent security holes,” and those do not show up in the same feedback loop that builds trust elsewhere.

Enforcing the tiers technically, not just as a policy

A rollout order written down and never enforced is a suggestion, not a control. If your tooling supports it, back the policy with actual restrictions on which paths an agent's tool calls can touch. Our guide to restricting an AI coding agent's file system access covers how to make tier 4 genuinely off-limits rather than just discouraged.

FAQ

Does this apply to a solo developer, or only teams?

It applies to both. A solo developer skips the social trust-building part but still benefits from deliberately sequencing exposure, since the actual risk (an agent making an expensive mistake before you have calibrated how much to trust it) is the same either way.

How long does tier 1 usually take?

Anywhere from a few days to a couple of weeks of regular use, depending on how much you are actually running it. The point is calibration through observation, not a fixed timeline.

What if the agent performs well immediately in tier 1?

Good performance on low-stakes work is evidence, but it is not evidence about high-stakes work specifically. Move up a tier at a time regardless, since the failure modes genuinely differ between tiers.

More on working safely with autonomous agents lives in our AI coding tools hub.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.