Dashboard

Restrict an AI Coding Agent's File System Access

A practical, tested setup for limiting where an AI coding agent can read and write: scoped directories, read-only mounts, ephemeral containers, git worktrees, and a real escape test.

Steve Jefferson
Steve Jefferson
Developer Advocate
12 September 20261 min read

Restrict an AI Coding Agent's File System Access

To restrict an AI coding agent's file system access, do not rely on a single permission prompt or a config flag. Combine four controls: run the agent from a scoped working directory that contains nothing but the repo it needs, mount everything else on the host read-only or not at all, run each task inside an ephemeral container or VM that gets destroyed when the task ends, and put the agent's changes in a separate git worktree so it never touches your main checkout. Each control alone is weak. Stacked together, a runaway edit, a bad shell command, or a genuinely malicious prompt injection has nowhere useful to go.

Why file access is the risk that actually bites

Most public writing about agent sandboxing focuses on network egress: can the agent phone home, exfiltrate a secret, or pull down a malicious script. That matters, but the more common failure in day to day coding is duller and more expensive: an agent given a broad working directory edits or deletes a file that was never part of its task. A sibling project on the same disk. A global git config. A credentials file it happened to have read access to because nobody scoped the directory tree it was launched from.

Coding agents run shell commands, write files, and often have permission to skip confirmation for actions the harness considers routine. If the routine boundary is "the whole home directory" instead of "this one repo," the agent's mistakes and a prompt injection's intentions look identical from the file system's point of view. Fixing this is a matter of shrinking what the agent can physically reach, not trusting it to stay inside lines you only asked it to respect.

Four controls that actually restrict file access

1. Scope the working directory to one repo, nothing above it

Never launch an agent from your home directory or from a parent folder that also contains other client repositories, SSH keys, or cloud credential files. Launch it from inside the repo, and make sure the repo's parent directory holds nothing else the agent could stumble into with a careless "../" in a shell command.

  • Bad: agent launched from ~/ with dozens of unrelated projects, dotfiles, and ~/.ssh visible one level up.

  • Good: agent launched from ~/work/agent-tasks/repo-name, a directory tree that holds only that checkout.

2. Read-only mount everything except the target repo

When the agent runs in a container, make the container's root file system read-only by default and mount only the working directory as read-write. This is a one-line flag in Docker and it turns "the agent could theoretically write anywhere" into "the agent can write to exactly one path, verified by the kernel, not by the agent's own judgment."

3. Run each task in an ephemeral container or VM

Spin up a fresh container from a clean base image per task, and destroy it when the task finishes. This does two things a long-lived dev container does not: it stops state and stray files from one task leaking into the next, and it caps the blast radius of any single bad run to whatever happened inside that one throwaway environment.

4. Isolate with a git worktree so the agent never touches your main checkout

A git worktree gives you a second working directory backed by the same repository, on its own branch, without cloning the repo again. Point the agent at the worktree, mount only the worktree path into its container, and your primary checkout (the one with your uncommitted work, your build artifacts, your IDE state) is on a different path the agent's container never sees. If the agent goes off the rails, it can wreck the worktree. It cannot touch the checkout you are actually working in.

Worked example: sandboxing one task end to end

Here is the sequence combining all four controls for a single agent task, using Docker and git worktrees. Adapt the image and flags to your own agent harness.

bash
# 1. Create an isolated worktree for this task only
git worktree add ../repo-agent-task-42 -b agent/task-42

# 2. Run the agent in a throwaway container:
#    --read-only locks the root filesystem
#    --tmpfs gives it a writable scratch dir that vanishes on exit
#    the only writable bind mount is the worktree itself
#    --network none removes exfiltration paths entirely for offline-capable tasks
docker run --rm -it \
  --read-only \
  --tmpfs /tmp \
  -v "$(pwd)/../repo-agent-task-42:/workspace:rw" \
  -w /workspace \
  --network none \
  --cap-drop ALL \
  agent-image:latest

# 3. Review the diff before anything touches your main branch
cd ../repo-agent-task-42
git diff main...HEAD

# 4. Tear the environment down completely
cd ..
git worktree remove repo-agent-task-42 --force

Nothing here is exotic. It is a bind mount, a read-only flag, a worktree, and a container that gets deleted. The point is that these four ordinary pieces, used together, remove the need to trust the agent's judgment about where it should and should not write.

Verify the sandbox actually holds: the escape test

A sandbox you have not tried to break is a sandbox you are guessing about. Before trusting the setup above with a real task, deliberately try to escape it from inside the container, either by hand or by asking the agent to run each command and report what happened.

Test command

Expected result

What it catches

touch /etc/canary.txt

Permission denied (read-only root)

Root filesystem is not actually read-only

cat ~/.ssh/id_rsa or /root/.aws/credentials

No such file or directory

Credentials leaked into the container image or mount

touch ../outside-worktree.txt

Permission denied or path not found

Bind mount scope is wider than the single worktree

git worktree list (from inside the container)

Only this task's worktree is visible

.git commondir exposes sibling worktrees or the main checkout

curl or ping any external host

Network unreachable

"--network none" was dropped or overridden somewhere in the run script

If any of these succeed when they should fail, the sandbox does not hold yet. Fix the specific control that leaked and rerun the whole table, not just the one test that failed. A sandbox test is only useful as a full suite, run every time you change the container image, the run script, or the harness.

Mistakes that quietly break an otherwise good sandbox

  • Mounting the repo's parent directory instead of just the worktree, so the agent can walk up and see the main checkout through git's commondir.

  • Dropping --read-only or --network none temporarily to fix one broken build step, then forgetting to restore it before the next task.

  • Reusing one long-lived container across many tasks for speed, which quietly turns an ephemeral sandbox into a persistent one with accumulated state.

  • Running the container as root without dropping capabilities, so a container-level exploit has far more to work with than the task ever needed.

  • Leaving cloud credential files or SSH keys inside the image or the mounted path because they were convenient during setup and never removed.

Frequently asked questions

Does a git worktree by itself sandbox an AI coding agent?

No. A worktree isolates which branch and commits the agent can affect, but it does not stop the agent from reading or writing files outside the repository if it is not also run inside a container or VM with restricted mounts. Worktree isolation is one of the four controls, not a replacement for the others.

Does read-only mounting slow down an AI coding agent?

Not meaningfully. Read-only mounts are enforced by the kernel with no measurable overhead. The agent only notices when it tries to write somewhere it should not be writing anyway, which is the entire point.

What is the fastest way to check if a sandbox actually works?

Run the escape test table above before the first real task and after any change to the container image or run script. It takes a few minutes and directly answers whether the restrictions are enforced or just assumed.

Should every AI coding agent task get its own container?

For anything beyond quick, supervised edits, yes. A fresh container per task keeps state from one run out of the next and limits how much damage a single bad run can cause, since the whole environment is discarded when the task ends.

Is restricting file access enough, or does network access matter too?

File access controls the blast radius of a mistake on disk, but they do not stop data leaving the machine. For tasks where that matters, pair file system restrictions with network-level sandboxing, which covers outbound network reach specifically.

The controls above exist because AI coding agents can now generate more changes than a team can manually review, which is also why CI infrastructure itself has become a funding thesis: see our coverage of Blacksmith's $45 million Series B for why validating AI-generated code is turning into a compute bottleneck.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.