Which Parts of Your Codebase to Let an AI Agent Touch First
A staged rollout order for giving an AI coding agent access to your codebase, starting with what is safe to get wrong and ending with what is not.
Which Parts of Your Codebase to Let an AI Agent Touch First
Give a new AI coding agent access in the same order you would give a new hire access: start where a mistake is cheap and visible, and only extend reach once it has earned it. In practice that means tests and internal tooling first, then well-covered feature code, then anything touching money, auth, or data migrations, and only once you have watched it work on the earlier tiers.
The rollout order, and why it is in this order
Tier | Example areas | Why this tier, why this order |
|---|---|---|
1. Safe to get wrong | Test files, internal scripts, documentation, non-customer-facing tooling | Mistakes here are caught by CI or a human before they ship anywhere, and you learn the agent's habits at zero cost |
2. Well-covered feature code | UI components, API endpoints with existing test coverage | Coverage acts as a safety net, so you are watching for review-worthy issues rather than catching untested regressions |
3. Under-covered feature code | Older modules, code nobody has touched in a while | Worth pairing with a request to write tests first, since this is where an agent's confidence and your test coverage are both weakest |
4. Sensitive by default | Auth, payments, permissions, data migrations | Errors here are expensive, sometimes irreversible, and often invisible until a customer hits them |
What “earning it” actually looks like
There is no fixed number of pull requests that unlocks the next tier. What you are actually watching for is whether the agent's mistakes in the current tier are the kind you expected (small, caught by tests, easy to explain) rather than the kind that surprise you (silent scope creep, tests it wrote to pass rather than to verify, confidently wrong reasoning about why something works). If mistakes are getting more surprising rather than less as you extend its reach, that is a reason to hold at the current tier, not push forward on schedule.
This is different from asking how much it should see
A related but separate question is how much of the codebase an agent needs in context to work well, which is mostly a technical constraint about window size and retrieval quality. Our piece on how much of your codebase an AI coding agent should see covers that. This piece is about a policy decision: given that it can see and edit anything you point it at, what should you actually authorize it to touch, and in what order, regardless of how much context it technically has access to.
The sensitive tier deserves its own rule, not a graduation
Auth code specifically is worth treating as a standing exception rather than something that becomes fair game once trust is established elsewhere. Our dedicated piece on letting an AI coding agent touch your auth code goes into why this tier resists the same earned-trust logic that works for everything above it: the failure mode is not “more bugs,” it is “silent security holes,” and those do not show up in the same feedback loop that builds trust elsewhere.
Enforcing the tiers technically, not just as a policy
A rollout order written down and never enforced is a suggestion, not a control. If your tooling supports it, back the policy with actual restrictions on which paths an agent's tool calls can touch. Our guide to restricting an AI coding agent's file system access covers how to make tier 4 genuinely off-limits rather than just discouraged.
FAQ
Does this apply to a solo developer, or only teams?
It applies to both. A solo developer skips the social trust-building part but still benefits from deliberately sequencing exposure, since the actual risk (an agent making an expensive mistake before you have calibrated how much to trust it) is the same either way.
How long does tier 1 usually take?
Anywhere from a few days to a couple of weeks of regular use, depending on how much you are actually running it. The point is calibration through observation, not a fixed timeline.
What if the agent performs well immediately in tier 1?
Good performance on low-stakes work is evidence, but it is not evidence about high-stakes work specifically. Move up a tier at a time regardless, since the failure modes genuinely differ between tiers.
More on working safely with autonomous agents lives in our AI coding tools hub.
How did this land?
About the author

Staff Engineer, Platform
Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.


