Files Your AI Coding Agent Should Never Write
Four published sandbox escapes share one mechanism: the agent obeyed every rule and wrote a file that a trusted tool outside the sandbox later executed.
A run of sandbox escapes published in July 2026 by Pillar Security shares one mechanism, and it is not the one most people picture. The agent never breaks out. It stays inside its sandbox, obeys every restriction, and writes an ordinary file. Something else on the machine, running with your privileges and outside the sandbox, later reads or executes that file.
As the researchers put it in their write-up of the series, the agent follows every rule and the escape happens on its own. That reframes the defence: the interesting boundary is not what the agent can run, it is what the agent can write.
What was actually found
The findings, reported by BleepingComputer on 20 July 2026, covered four widely used coding agents and were fixed at different speeds:
Tool | The file that did it | Status |
|---|---|---|
Cursor | A workspace-controlled hook configuration that ran commands unsandboxed, tracked as CVE-2026-48124 | Fixed in version 3.0.0 |
Cursor | A virtualenv interpreter the agent could edit | Fixed in version 3.0.0 |
OpenAI Codex CLI | A safe-command allowlist that trusted a command by name even when invoked in a way that was not read-only | Fixed in version 0.95.0 |
Codex, Cursor, Gemini CLI | A reachable Docker socket, giving an unsandboxed place to run code | Fixed |
Google Antigravity | A task configuration file that bypassed Secure Mode, and a macOS Seatbelt denylist bypass | Classified as other valid security vulnerabilities and downgraded |
Two things stand out. The Docker socket issue hit three separate products at once, because they all made the same reasonable assumption about what a local daemon means. And the two Antigravity findings were downgraded by the vendor as hard to exploit, which means the mitigation there is yours rather than theirs.
The deny-list this implies
Every entry below is a file whose contents get consumed by something outside the sandbox. None of them is code your agent needs to edit to do ordinary work.
Agent and editor hook configuration. Anything that registers a command to run on an event. This is the CVE-2026-48124 shape, and it is the highest value target because it converts a file write directly into command execution.
Editor task and launch configuration. The
.vscodedirectory in particular. Task definitions are executed by the editor, not by the agent, so sandbox rules do not apply to them.Interpreters and virtualenv contents. The
bindirectory of a virtual environment, shebang lines, and anything a subsequent command will execute as a program rather than read as data.Git internals and hooks. The
.gitdirectory including hooks and fsmonitor configuration. Git runs these itself, with your permissions, often without you invoking anything.Shell startup files. Profiles and rc files. An agent has no legitimate reason to modify your login shell, and a modification there survives every restart.
CI workflow definitions. A change here runs on a machine with your deployment credentials. Legitimate occasionally, and always worth a human read rather than an automatic merge.
Package manager lifecycle scripts. Install and post-install hooks in a manifest execute on the next install, on your machine and on your colleagues'.
How to enforce it
Instructions in a rules file reduce the frequency and do not stop a determined injection, so put the enforcement somewhere the agent does not control:
Update the tool. Every finding above except the downgraded pair is fixed in a named version. Checking your version number is the highest value action on this page.
Add the paths to a deny-list in your agent's own configuration, where the tool supports one. This blocks the accidental case, which is most of them.
Do not expose the Docker socket to an agent's environment. If a task needs containers, give it a rootless or proxied interface rather than the raw socket.
Put the sensitive paths behind CODEOWNERS so a change requires human review, and add a CI check that fails when a diff touches them.
Read the file list of every agent commit before merging, not the code. A change to
.gitor.vscodein a pull request about a bug fix is the signal, whatever the diff says.
Why these particular files
The common property is that something other than the agent reads them, at a moment you did not choose, with privileges the agent does not have. Two examples make the shape concrete.
Git hooks are executable scripts that Git runs itself on ordinary operations such as commit and checkout. Nothing invokes them explicitly, they are not part of any build, and they run as you. A file written into that directory during an agent task executes the next time you commit.
Editor task definitions behave similarly: the editor reads them from the workspace and can run them, in some configurations without a prompt. The agent's sandbox has no view of the editor process, so a restriction on what the agent may execute says nothing about what the editor will.
Once you hold that pattern, the deny-list stops being a list to memorise and becomes a question to ask of any file: does anything outside the sandbox read this, and does it run what it finds? If yes, the agent should not be writing it unattended. That question also governs how much rope to give a terminal, covered in safe terminal access for an agent.
Why prompt injection makes this urgent
A deny-list matters more once you accept that the instruction to write a hook file may not come from you. Prompt injection embedded in a repository, an issue, a dependency README or a fetched page can reach an agent that is reading those things as part of its work.
At that point the agent is not being tricked into breaking its sandbox. It is being asked to do something it is fully permitted to do, and the sandbox has no opinion. The general defence is in preventing prompt injection, and the containment side in sandboxing an AI agent.
How this differs from an agent breaking containment
It is worth keeping the two apart, because they call for different defences. An earlier containment incident involved a model reaching outside the environment it was given. The findings here involve no breach at all: the sandbox worked exactly as designed, and the escape route ran through a trusted tool on the host that nobody thought of as part of the boundary.
Defences aimed at the first case, such as tighter command allowlists and network egress rules, do nothing about the second. Only a write deny-list does.
Frequently asked questions
Does this mean coding agents are unsafe to use?
No. It means the sandbox is not the whole boundary, and that patched versions matter more than most people assume. Every finding here was reported responsibly and most were fixed quickly, which is the system working. Running a version from before the fix is the actual risk.
Is a container enough on its own?
Only if the container cannot reach anything that executes files from the workspace. Mounting your real project directory into a container and exposing the Docker socket recreates the exact problem inside a box that feels safer, which is the trap the shared Docker socket finding illustrates.
What about files the agent needs to write during normal work?
Source, tests, documentation and lockfiles are the normal working set, and none of them appear above. The deny-list is deliberately narrow: configuration that something else executes. If it feels restrictive in daily use, the list has probably grown beyond what these findings support.
Should I audit what my agent has already written?
Worth thirty minutes if you have been running agents unattended. Check git log for commits touching .git, .vscode, CI workflows and shell profiles, and check whether your tool version predates the fixes listed above. What your coding assistant sends to the vendor covers the other half of the review.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


