Dashboard

How Much of Your Codebase Should an AI Agent See?

More context is not better context. Dilution, pattern contamination and scope creep all get worse as you add files. A narrow slice plus an accurate map wins.

Steve Jefferson
Steve Jefferson
Developer Advocate
17 September 20261 min read

How Much of Your Codebase Should an AI Agent See?

An AI coding agent should see the files it needs to change, the files those directly depend on, and a written map of everything else. Not the whole repository. The instinct to give an agent maximum context comes from a reasonable place, since a human joining your team would benefit from reading everything, but it produces worse results in practice: slower responses, higher cost, and a measurably greater tendency to modify things you did not ask it to touch.

Why more context makes output worse

Three separate effects stack up, and only the first is widely discussed.

  1. Dilution. Attention spreads across everything in the window. In the transformer architecture every token attends to every other token, so an instruction competes with fifty thousand lines of unrelated code for the model's attention. The instruction does not get stronger when you add more around it.

  2. Pattern contamination. If your repository contains an old module written in a style you abandoned, an agent that can see it will treat it as a valid example of house style, and you will get new code written in the style you were trying to leave behind.

  3. Scope creep. An agent that can see an adjacent bug will often fix it, uninvited, in the same change. This feels helpful until you are reviewing a diff that touches nine files when you asked for one.

There is also the plain cost of it. Every file in context is re-read on each turn, which is the same mechanism that makes long conversations slower and the reason a large working set gets expensive quickly.

What an AI coding agent should actually see

A working default for a single task, in priority order:

Include

Why

The files being changed

Obvious, and often the only thing people include

Direct dependencies and callers

Prevents signature changes that break unseen call sites

One good example of the pattern to follow

Far more effective than describing the pattern in words

The relevant test file

Gives the agent a way to verify itself

A conventions file

Folder layout, error handling, data fetching, naming

Type definitions or schema for the data involved

Removes the most common category of invented field names

The conventions file and the single good example are the two that punch above their weight. Together they do most of the work of making an agent follow your code style, and they cost nothing per task once written.

What to exclude, and how

Some exclusions are about quality. Others are about not leaking things.

  • Secrets, environment files and anything with a credential in it. This is non-negotiable and it is the reason to configure exclusions rather than rely on remembering.

  • Build output, lock files, minified bundles, vendored dependencies and generated code. Enormous, and it teaches the agent nothing.

  • Customer data, fixtures with real personal data, and database dumps. Anonymised fixtures only.

  • Deprecated modules and code you are actively migrating away from, unless that migration is the task.

  • Large binary assets, which waste the window without contributing anything readable.

Most agent tools support an ignore file, typically following the same pattern syntax as your version control ignore rules. Configure it once per repository and check it into the repository so it applies to everyone. Relying on each person's memory fails the first time someone is in a hurry, and the failure mode for secrets is the expensive one: recovering from an agent committing a secret to your repo is far more work than writing the ignore rule.

The map beats the territory

The single highest-leverage thing you can give an agent is not more code, it is a short architecture document: what the main modules are, what each is responsible for, where things live, and which boundaries must not be crossed. Two hundred words of accurate map is worth more than twenty thousand lines of source, because it answers the question the agent would otherwise try to infer by reading everything.

Write it once, keep it at the repository root, and update it when the structure changes. If you do not have one, an agent can draft it for you: ask it to explain the codebase and then correct what it got wrong. Your corrections are the document, and they are exactly the knowledge the agent was missing.

Scaling it by repository size

The right answer changes with size, and the transition points are fairly sharp.

  1. Under a few thousand lines: whole repository is fine, and the overhead of being selective is not worth it.

  2. Ten to a hundred thousand lines: selective context per task, plus a conventions file and an architecture map. This is where most people are and where most of the benefit lives.

  3. Larger than that, or a monorepo: scope the agent to one package or service at a time, with the interfaces of the packages it talks to. Treat crossing a package boundary as a separate task rather than a bigger one.

FAQ

Should I let an AI agent index my entire repository?

Indexing for retrieval is different from putting everything in the context window, and it is generally fine: retrieval pulls in relevant pieces on demand. The rules about secrets and customer data still apply, because indexing means a copy of that content leaves your machine.

Does a bigger context window change this answer?

Only at the margins. A bigger window raises the ceiling on what fits, but dilution and scope creep are properties of what you include rather than of what fits, which is broadly why a bigger context window does not mean better answers.

What if the agent keeps missing files it needs?

That is usually a map problem rather than a context-size problem. An agent that knows where things live will ask for the right file. One that is guessing will either ask for everything or invent what it cannot find.

Is it safe to give an agent read access to production code?

Read access to source is normally fine once secrets are excluded. Access to production systems and production data is a separate decision with a much higher bar, and the two get conflated more often than they should.

A related but separate question is the order in which to extend that access as trust builds, covered in which parts of your codebase to let an AI agent touch first.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.