Dashboard

What Is a Confused Deputy Attack in AI?

A confused deputy attack tricks a program that holds real authority into using it on someone else's behalf. Prompt injection is how the request gets in. This is why it works.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
16 September 20261 min read

What Is a Confused Deputy Attack in AI?

A confused deputy attack tricks a program that holds real authority into using that authority on someone else's behalf. The program is not compromised and no credential is stolen. It is simply asked to do something, it has the permission to do it, and it cannot tell that the request came from somewhere it should not have.

This is the mechanism underneath most AI agent security incidents. Prompt injection is how the request gets in. Confused deputy is why the request works.

The original case, from 1988

The term comes from a paper by Norm Hardy describing a Fortran compiler called FORT, installed in a privileged system directory called SYSX. Because of where it lived, FORT could write to every file in that directory, including SYSX/BILL, the system's billing records.

FORT accepted a command-line argument telling it where to write debug output. A user passed SYSX/BILL as that argument. FORT wrote debug output over the billing file.

Nothing was hacked. FORT had legitimate write access, the user had no such access, and FORT could not distinguish "the operator wants debug output here" from "a user wants the billing file destroyed". It was the deputy, and it was confused.

Hardy's case is worth knowing because the shape repeats exactly, thirty-eight years later, in systems that share none of its technology.

The same bug, with an AI agent

Replace the compiler with an agent and the structure is identical.

An agent connected to your email has a token that can read and send mail. It reads an incoming message. That message contains text instructing it to forward the last ten emails to an external address. The agent has permission to forward email. It does.

The attacker never had your credentials. They sent you an email. The agent supplied the authority, and it could not distinguish an instruction from you from text that arrived inside data it was asked to process.

That last sentence is the whole problem. An agent receives your instructions and the content it works on through the same channel, as text, with nothing structurally marking which is which. Every tool the agent holds is a capability an attacker can try to borrow by getting text in front of it.

This is why prompt injection is not simply a content-filtering problem. Filtering the text is defence in depth. The underlying issue is that the agent's authority is ambient: it applies to whatever the agent decides to do, regardless of who prompted the decision.

Why permission checks do not save you

The intuitive fix is to check permissions harder. It does not work, and understanding why is the useful part.

When the agent forwards that email, every permission check passes. The token is valid. The agent is authorised to send mail. The action is within scope. An access control list answers the question "may this program do this?" and the answer is a truthful yes.

The question that needed asking was "on whose behalf, and does *that* party have this right?" Access-control-list systems do not carry that information at the point of the check. The authority comes from who the deputy is, not from who asked.

Capability-based systems address this by bundling the designation of a resource with the permission to use it, so a request arrives carrying its own authority rather than borrowing the deputy's. In practice, for people building on top of AI agents today, the applicable lessons are narrower but real.

What actually reduces the risk

Scope the token to the task, not to the agent. An agent that summarises your inbox needs read access. It does not need send. The most common version of this failure is granting a broad scope once, at setup, because it was easier than working out the minimum. That habit has a name and a cost: OAuth scope creep in AI agents.

Give the agent its own identity. When the agent acts under your account, every log entry says you did it and every permission you hold is available to it. A separate account bounds the damage and makes the audit trail readable. More on that in should an AI agent have its own user account.

Separate reading untrusted content from acting with authority. This is the highest-value structural change available. One component ingests the email, the web page, the support ticket. A different component, which never sees raw untrusted text, holds the credentials and executes actions against a constrained set of operations. The first can be fooled; it cannot do anything.

Require confirmation for irreversible actions, with the action described. A confirmation that says "allow this agent to send email?" is useless. One that says "send to attacker@example.com, subject: Fwd: Q3 contracts" is a control, because the confused deputy's action looks wrong the moment it is spelled out.

Constrain the destination, not just the operation. FORT's bug was that the output path was attacker-controlled. Most agent equivalents are the same: the operation is fine, the target is not. Allowlisting recipients, domains, or paths kills a large class of this.

For the implementation side of these, preventing prompt injection in your AI app and sandboxing an AI agent go deeper than the summary above.

Where this shows up in practice

Setting

The deputy

The confusion

Email agent

Agent holding a mail token

Instructions inside a received message

Browsing agent

Agent with your session cookies

Instructions inside a fetched page

Support triage

Agent with CRM write access

Instructions inside a customer ticket

Coding agent

Agent with repo and CI access

Instructions inside a dependency's README or an issue comment

MCP server

Server holding a long-lived API key

A request relayed from a caller whose rights were never checked

The last row is the one growing fastest. An MCP server that holds a credential and executes whatever a connected model asks is, structurally, FORT with a JSON interface. Whether that is a problem depends entirely on whether anything checks what the calling context was actually entitled to.

Why the name is worth using

Most writing about agent security stops at "prompt injection is dangerous", which is true and not actionable. Naming the confused deputy gives you the design question: which components hold authority, and can they tell who is asking?

That question has answers. "Do not let the model be tricked" does not. For the wider risk landscape, see the AI risks guide.

FAQ

What is a confused deputy attack in simple terms?

A program with permissions is tricked into using them for someone who does not have those permissions. The program is not hacked. It is fooled about who is asking.

How is it different from prompt injection?

Prompt injection is the delivery method, getting malicious instructions into what the model reads. Confused deputy is why those instructions have effect: the agent holds real authority and cannot tell the instruction did not come from you.

Where does the term come from?

A 1988 paper by Norm Hardy, describing a Fortran compiler with write access to a privileged directory that was asked to write its debug output over the system billing file.

Do permission checks prevent it?

No. Every check passes, because the deputy genuinely does have the permission. What is missing is any record of who the request is really on behalf of.

What is the single most effective defence?

Separating the component that reads untrusted content from the component that holds credentials and acts. The first can be fooled without being able to do anything.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.