Dashboard

How to Limit What an AI Agent Can Reach on Your Network

Open network access is the default for most agent setups and the wrong one. Here is how to build an allowlist that survives a persistent agent.

Steve Jefferson
Steve Jefferson
Developer Advocate
24 September 20261 min read

The control that matters is not a rule in your prompt. It is a proxy the agent has to go through, with a list of hosts it is allowed to reach and a denial for everything else. Prompt instructions are requests. An egress allowlist is a wall, and it works whether the agent is being manipulated, is confused, or is simply trying very hard to finish a task.

That last case is the one people underestimate. Research published this week found agents escalating to exploit attempts against public data sites after ordinary retrieval failed, with nobody asking them to, per the Transluce write-up. Persistence is a feature you tuned in deliberately. Network scope is how you bound it.

Start by writing down what it actually needs

Before any config, list the hosts. Most agents need a shorter list than their owners expect:

  • The model API endpoint.

  • Your own application or database host, often via an internal name.

  • Package registries, but only if the agent installs things at runtime, which it usually should not.

  • Whatever specific third-party APIs the job requires, named individually.

If the answer is "it needs to browse the web", that is a different risk class and deserves its own decision, made explicitly rather than by omission. Everything else is a closed list, and closed lists are easy to enforce.

Put the control at the network edge

The pattern that holds up is: agent runs in its own container, container has no direct route out, all traffic goes through an HTTP proxy that enforces the allowlist. The agent cannot opt out because there is no other path.

A minimal Squid configuration expresses the whole idea:

acl allowed_hosts dstdomain api.anthropic.com
acl allowed_hosts dstdomain .internal.example.com
acl allowed_hosts dstdomain api.stripe.com

http_access allow allowed_hosts
http_access deny all

# refuse CONNECT to anything but 443, so the proxy is not a tunnel
acl SSL_ports port 443
http_access deny CONNECT !SSL_ports

Then run the agent's container with no default route and HTTPS_PROXY pointed at that proxy. In Docker terms that is an internal network for the agent plus a second interface on the proxy. In Kubernetes it is a NetworkPolicy with an empty egress rule set apart from the proxy, which is more reliable than trying to enumerate IP ranges.

The reason to do this at the proxy rather than with firewall rules is that modern APIs sit behind CDNs with address ranges that change without notice. Names are stable, addresses are not. Filter on names.

The two holes almost every allowlist has

**Link-local metadata endpoints.** On every major cloud, 169.254.169.254 serves instance credentials to anything that can make an HTTP request from the host. An agent that can reach it can read the role credentials of the machine it runs on, and from there your allowlist is irrelevant because it has real keys. Block the entire 169.254.0.0/16 range at the network level, not just in the proxy, and confirm the block from inside the container rather than assuming the platform did it.

**Names that resolve inward.** An allowlist of domain names is only as good as what those names resolve to. internal-tools.example.com pointing at 10.0.3.14 is fine if you meant it and a hole if you did not, and an attacker-controlled domain can resolve to a private address deliberately. The fix is to reject responses that resolve into private ranges unless the host is explicitly permitted to, which most proxies support as a DNS or destination-address rule. This is server-side request forgery wearing agent clothing, and it is the bug class that outlives every other item on this list.

Cap the failures too

Scope limits where an agent can go. It does not limit how hard it tries, and repeated failure is where improvisation begins. Two limits are worth setting alongside the allowlist:

  1. A retry ceiling per task, low enough that a wall stays a wall. Three is usually plenty.

  2. An explicit, blessed failure path. An agent with permission to return "I could not retrieve this" does not need to invent a route around the obstacle. If your prompt implies the task must be completed, you have removed the option of giving up.

What this looks like in the three common setups

The principle is identical. The mechanism differs enough to be worth naming.

**Agent in a container on a server.** The clean case, and the one described above. Internal Docker network for the agent, proxy on a second interface, no default route for the agent container. Verify by trying to reach an unlisted host from inside, since an internal network that is not actually internal is the usual mistake.

**Agent in Kubernetes.** Use a NetworkPolicy that permits egress only to the proxy service and to kube-dns, and nothing else. Resist enumerating allowed CIDRs in the policy itself: cloud service ranges change, the policy silently rots, and the failure mode is an outage rather than a leak, so nobody discovers the rot until it breaks something at an inconvenient moment. Note that NetworkPolicy needs a CNI that enforces it, and a cluster where nothing enforces policies will accept your manifest and ignore it.

**Agent on a developer laptop.** The hardest to lock down and the most common place agents actually run. You cannot easily restrict egress on a machine the developer controls, and you should not try. Run the agent inside a container with the proxy configuration instead, which bounds both network reach and filesystem reach in one move, and treat any credentials the agent can see as credentials it may transmit.

For anything running against production data, the container route is not optional regardless of shape. An agent on a laptop with direct database credentials and open egress is the configuration behind most of the incidents worth avoiding.

Log every request, not just every answer

Most teams log what the agent concluded. Far fewer log every outbound request it made, which is the record you need when someone else's abuse report arrives. Your proxy already sees all of it, so turn on access logging and keep destination, timestamp and agent run identifier. The cost is negligible and the alternative is learning what your agent did from a stranger.

Pair it with a simple alert on denied requests. A handful is normal. A burst of denials to hosts nobody configured is the signature of an agent working a problem, and it is the earliest signal you will get.

Verify from inside

An allowlist you have not tested is a hypothesis. Exec into the running container and check the four cases:

curl -sS -o /dev/null -w '%{http_code}
' https://api.anthropic.com/v1/models   # expect a response
curl -sS --max-time 5 https://example.com                                        # expect denied
curl -sS --max-time 5 http://169.254.169.254/latest/meta-data/                   # expect no route
getent hosts internal-tools.example.com                                          # check what it resolves to

Run this after every infrastructure change. Allowlists rot quietly, usually when somebody adds a convenience rule during an incident and nobody removes it.

Where this fits with the rest

Network scope is one of three boundaries, and it is the one that holds when the others fail. Credentials are the second: an agent with read-only database access cannot write, regardless of where it can connect, and giving a coding agent read-only production access covers that side. Tool permissions are the third, and vetting an MCP server before connecting it applies before you hand an agent a new capability at all.

Prompt instructions are a fourth layer and the weakest one. Keep them, since they reduce the volume of attempts. Do not count them.

Questions

Is a prompt instruction not enough?

No. An instruction shapes behaviour but does not constrain capability, and it fails exactly when you need it: under injection, under confusion, and under a model that is determined to finish the job.

What about agents that legitimately need to browse the web?

Treat open browsing as a separate, explicit decision rather than a side effect of open egress. Route it through the same proxy with a broader policy, keep the private-range and metadata blocks absolutely, and log everything. The risks specific to that mode are in whether it is safe to let an AI agent browse the web.

Does this apply to a coding agent on my laptop?

The metadata endpoint risk is lower, everything else applies. The practical local equivalent is running the agent in a container with a proxy rather than directly on your machine, which also bounds what it can read from your filesystem.

How do I allow a service whose addresses change constantly?

Allow the domain name, not the address, and let the proxy resolve at request time. That is the main reason to filter at an HTTP proxy rather than with address-based firewall rules.

What is the single highest-value thing to do first?

Block the link-local metadata range and confirm the block from inside the container. It takes a few minutes and closes the path from "agent made an odd request" to "agent has your cloud credentials".

For the wider set of failure modes worth designing against, AI risks: a practical guide for builders covers the territory.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.