AI Agent Incident Reporting: Inside SAFE

The Linux Foundation wants AI agent incidents reported on aviation-style deadlines. Here is what the SAFE draft asks for, and what it leaves out.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
14 August 20261 min read

On 4 August 2026 the Linux Foundation published a request for comments on the Shared AI Findings Exchange, or SAFE: a proposal that organisations running AI systems report security incidents involving those systems on fixed deadlines, to a central exchange that would look for patterns across reports and publish recommendations. It is a draft, not a rule, and nobody is bound by it yet.

It is still worth reading, because it is the first serious attempt to define what an AI agent incident even is.

Where the idea comes from

Justin Boitano of Nvidia has described the model as based on NASA's Aviation Safety Reporting System: a scheme where operators report near misses as well as accidents, the reports are pooled, and the value comes from the pattern rather than any single filing.

The alliance behind it, the Open Secure AI Alliance, counts over a hundred member organisations. Cisco, CrowdStrike, Hugging Face, Nvidia and Red Hat are named as drafting partners in the Linux Foundation's request for comments.

What counts as reportable

The draft is more specific than most AI safety documents. Reportable events include:

  • An AI system gaining unauthorised access to, or exploiting, a third-party system

  • An AI system disclosing confidential information

  • An agent continuing to probe a production target after its operators suspect the activity is unauthorised

  • Near misses, meaning the same events where nothing was ultimately damaged

That last one is where the aviation analogy earns its place. A near miss is exactly the category most organisations currently handle by closing the terminal and saying nothing.

The clock

Step

Deadline

Notify customers of credible data exposure

72 hours

Confidential initial report to the exchange

4 business days

Preliminary public factual report, where appropriate

30 days

Remediation update

90 days

Members would also preserve evidence: prompts, agent traces, tool calls, identities, permissions and credentials. All requirements are qualified as subject to security, legal and investigative constraints, which is doing a fair amount of work in a document this short.

If you cannot produce that evidence today, that is the actionable finding here regardless of whether SAFE ever ships. An agent whose tool calls are not logged cannot be investigated, by you or anyone else. Keeping an audit trail of AI use is the prerequisite, and it is not a large project.

The two gaps

The first is legal. The draft offers no formal safe-harbour protection. An organisation that files an honest report about its own agent exfiltrating customer data has produced a dated, written admission with regulatory and litigation value, and SAFE does not shield it. Aviation reporting works partly because reporters get protection. Without it, the reports that arrive will be the ones that were already public.

The second is who signed. OpenAI and Anthropic, the two vendors whose models sit underneath a large share of production agents, are not members. A findings exchange missing the two largest sources of findings is a smaller instrument than it looks.

How it differs from breach notification

If this feels familiar, it is because data breach notification regimes already impose similar clocks. The difference is what triggers them. Breach rules fire when personal data is exposed. SAFE fires when an AI system behaves outside its authorisation, whether or not any data moved.

That is a meaningfully wider net. An agent that discovered a path into a partner's staging environment and stopped there has exposed nothing and breached nothing. Under SAFE it is still a reportable near miss, because the pattern is the point rather than the damage.

It also means the two regimes can fire on the same incident with different deadlines and different recipients, which is an argument for one internal process that satisfies the strictest of them rather than two that partly overlap.

What to do with this now

Nothing about SAFE requires action from a small team. But the definitions are useful even as a draft, because they answer a question most incident response plans skip: what, specifically, is an AI incident?

Three things are worth borrowing:

  1. Treat near misses as incidents. The agent that tried to email a customer list and was stopped is data you will want in three months.

  2. Log the four artefacts. Prompts, tool calls, identities, permissions. Everything else in an investigation is reconstructable from those.

  3. Decide the notification question before it happens. Seventy-two hours is a plausible clock. Deciding who makes that call while the incident is running is how organisations miss it.

If you do not have an incident plan that mentions agents at all, start there. And if the underlying worry is what an agent can reach in the first place, the containment question comes before the reporting question: sandboxing an agent limits the size of the report you will one day have to write.

The comment period is open, and the draft will change. What will not change is that the question has been asked in public, with dates attached. That tends to be how voluntary frameworks become expected practice, and then eventually the thing an insurer asks about. Following where these proposals go next is cheaper than being surprised by them.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.