Dashboard

How to Tell If an AI Agent Has Been Compromised

No malware, no unusual login, no privilege escalation. The credentials being used are the ones you gave it. Detection lives in the sequence, not the actions.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
4 September 20261 min read

A compromised AI agent does not look compromised. It looks like an agent having an unusually productive afternoon. It reads files it has never read, calls an endpoint it has never called, and reports success, because whoever redirected it also told it what to say. Traditional intrusion signals do not fire: no malware, no unusual login, no privilege escalation. The credentials being used are the ones you gave it.

This is a detection problem, and it is solvable, but only if you decide in advance what normal looks like.

What compromise means here

Two mechanisms, one outcome.

The agent was redirected. Instructions reached it through content it processed: a support ticket, a scraped page, a code comment, a file it was asked to summarise. It is doing what it was told, by someone who is not you. This is prompt injection, and it is by far the more likely of the two. OWASP ranks it first in its top ten risks for LLM applications, which matches what turns up in practice.

The agent's credentials were taken. Its API key or token is being used by something else entirely. Rarer, and usually downstream of a leaked key rather than a redirected agent.

In both cases you see an authorised identity doing authorised things in an unauthorised sequence. That last word is where detection lives.

Six signals that an AI agent has been compromised

A tool call the agent has never made before. The single strongest signal. Agents in production are boringly repetitive: the same five or six tools, in similar proportions, day after day. A support triage agent that suddenly calls the outbound email tool is worth a page at three in the morning.

A change in the read to write ratio. Most agents read far more than they write. A sustained shift toward writes, especially deletes and updates, is either a compromise or a bug, and both warrant stopping.

Output that references something outside the task. The agent mentions a system, a person, or a file that has nothing to do with what it was asked. Injected instructions leak into responses more often than you would expect, because the model is trying to be helpful about both tasks.

A spike in token use per task. Injected instructions add work. A task that normally costs 4,000 tokens costing 30,000 means the model is doing something extra. This is the cheapest signal to instrument, since you are already logging it for billing.

Activity outside its window. Agents triggered by tickets, commits, or schedules have a shape to their day. Traffic at hours the trigger never fires is a strong signal, and it is trivial to alert on.

Refusals that stop happening. If your agent normally declines certain requests a few times a week and that count drops to zero, something changed. Silence is data.

Build the baseline before you need it

None of the above works without a two-week record of normal. Collect it now, while nothing is wrong.

For each agent, write down:

  • The complete list of tools it may call, and roughly how often each

  • Median and 95th percentile tokens per task

  • The hours and days it is normally active

  • The systems it touches, by name

  • What its output normally looks like, including length

That document is both your detection baseline and your permissions specification, and writing it usually reveals that an agent has access to three things it has never used. Remove those. Capability an agent does not need is capability an attacker inherits.

Instrumentation that pays for itself

Log the full trajectory, not just the result. Every model call, every tool invocation, every tool result. Post-incident, the answer to "when did it turn" is always in the trajectory and never in the final output. This is the same argument as keeping an audit trail of AI use.

Give the agent its own identity. An agent sharing a human's credentials is undetectable by definition, because every action looks like that person's. A dedicated service account makes anomalies visible and revocation instant, which is the case made in should an AI agent have its own user account.

Alert on tool-call novelty, not volume. Volume alerts are noisy and fire on busy days. First-time-ever tool calls are rare and almost always interesting.

Put a diff in front of destructive actions. Not every write needs approval. Deletes, permission changes, outbound messages, and anything touching money should show a human what is about to happen. A compromised agent will happily explain why the deletion is necessary; the diff does not care what it says.

When you think it happened

In order, and quickly.

  1. Revoke the agent's credentials. Not pause the process, revoke the token. A paused process with a live token is not contained.

  2. Freeze the trajectory logs. Copy them somewhere the agent cannot write. Retention windows are how evidence disappears.

  3. Find the entry point. Work backwards from the first anomalous action to the last untrusted content the agent processed. It is nearly always a document, a ticket, a page, or a file.

  4. List every action after that point. Assume all of them were directed by the attacker, including the ones that look fine.

  5. Check what it read, not only what it wrote. Exfiltration is quieter than damage and often the actual objective.

  6. Rotate anything it could see. Every credential in its environment, not just the one it used.

Then write it up. The structure for that is in building an AI incident response plan, and the version written before an incident is worth several written after one.

The honest limitation

You cannot reliably detect a compromise that stays inside the agent's normal behaviour. An injection that makes a code review agent approve one bad pull request produces no anomaly at all: expected tool, expected volume, expected hours.

Which is why detection is the second line. The first is containment: scope the permissions so that the worst available action is survivable. Keeping the agent inside a sandbox reduces the blast radius, and reducing blast radius is more tractable than perfect detection will ever be. The wider AI risk landscape sits in our guide to the risks worth planning for.

FAQ

Will my existing security tooling catch this?

Mostly not. Endpoint and network tools look for unauthorised access, and this is authorised access being misdirected. The signal lives in your application and agent logs.

How long do compromises usually go unnoticed?

Until something visibly breaks or someone reads the logs. Without trajectory logging and a baseline, there is no mechanism that would surface it, which is the real answer to the question.

Can I ask the agent whether it has been compromised?

No. If injected instructions are in its context, they can shape that answer too. Detection has to sit outside the model, in your logs and alerts.

Does a smaller or local model reduce the risk?

Not meaningfully. Injection targets the instruction-following behaviour every capable model has. Permissions and isolation reduce the risk; model choice barely moves it.

What is the single highest-value thing to add first?

Trajectory logging with a first-time tool-call alert. It is a small amount of work, it detects the most common pattern, and without it every other step in an investigation is guesswork.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.