Which AI Agent Actions Should Never Be Reversible
The useful question is not whether your agent will make a mistake. It is whether you can take the mistake back. Here is how to sort actions by reversibility.
Which AI Agent Actions Should Never Be Reversible
Sort your agent's actions by how hard they are to undo, not by how risky they feel. Anything you cannot reverse within a few minutes, using a mechanism you have actually tested, needs a human gate before it runs. Everything else can run freely and be cleaned up afterwards. That single distinction does more for you than any permission matrix.
The reason is simple: an agent that makes ten mistakes you can undo is fine. An agent that makes one mistake you cannot is a bad afternoon.
Reversibility is a spectrum, not a flag
Most permission systems ask whether an action is allowed. That is the wrong first question. Ask how expensive it is to take back.
Tier | Undo cost | Examples | Gate it? |
|---|---|---|---|
Free | Seconds, no trace | Read a file, run a query, draft text | No |
Cheap | Minutes, internal only | Create a record, write a file, open a branch | No |
Costly | Hours, needs coordination | Delete records, run a migration, change config | Usually |
One-way | Cannot be undone | Send an email, charge a card, post publicly, rotate a key | Always |
The tiers that matter are the bottom two, and the line between them is not about severity. Deleting a production table is severe and, if your backups work, recoverable. Sending one email to four hundred customers is less severe and completely permanent.
The four properties that make an action one-way
An action is effectively irreversible if any of these hold.
Someone else saw it. Email, SMS, a public post, a webhook you fired at a partner. The moment information leaves your boundary, you have lost the ability to unsend it, regardless of what the API's delete endpoint claims. A retracted Slack message was still read.
Money moved. Card charges, refunds, payouts, subscription changes. Reversing these is possible and expensive, involves a second system you do not control, and often leaves a record on someone's statement either way.
An external system committed. A DNS change that propagated, a package published to a registry, a commit pushed to a branch others have pulled, a third-party record created with an ID now referenced elsewhere.
The old state is gone. A key rotated without the old one stored, a file overwritten without version history, a record hard-deleted with no soft-delete column. These are the ones that hurt most, because they look reversible right up until you try.
Build the undo before you need the undo
The mistake teams make is deciding an action is reversible because it theoretically is. Reversibility you have not exercised is a hypothesis.
Three things make this concrete:
Soft delete everything the agent can delete. A deleted_at column turns a one-way action into a cheap one. If you are working in an app that was scaffolded quickly, adding soft delete and undo is usually a smaller change than it looks and buys you an entire tier.
Give outbound communication a delay. A five-minute hold queue on any email or message an agent sends converts the single most common one-way action into a recoverable one. Most teams that add this find they use the cancel button more than they expected.
Test the restore path quarterly. Not the backup. The restore. Run it, time it, write down the number. That number is what decides whether "delete records" sits in the costly tier or the one-way tier for you specifically.
Where the gate goes
A gate is a pause where a human confirms before the action runs. The failure mode is putting gates everywhere, because then people click through them without reading, and you have bought nothing.
Gate the one-way tier and nothing else. That usually means a short list: sends, charges, publishes, key rotations, hard deletes. If the list is longer than about eight action types, some of those actions belong in a lower tier and the right fix is to make them reversible rather than to gate them.
The pause is most useful when it shows the payload, not the intent. "Agent wants to send an email" is not reviewable. "Agent wants to send this subject line to these 412 addresses" is. Our guide on reviewing an agent's plan before it runs covers what a reviewable plan contains, and the same standard applies at the gate.
The special case of spending
Money is worth pulling out separately because it has a control that nothing else has: a hard ceiling. You cannot cap how embarrassing an email is, but you can cap how much an agent can spend before it stops. Setting spending limits for agents is the cheapest risk reduction available in this whole area, and it works even when your gating logic has a hole in it, because the limit lives in a different system.
If it has already gone wrong, what to do after an agent makes an unauthorized purchase walks the recovery.
A worked example
Take an agent that handles inbound support email. Its plausible action list:
Read the ticket, search the knowledge base, look up the customer. Free tier, no gate.
Draft a reply, add internal notes, tag the ticket. Cheap tier, no gate.
Issue a refund, delete the customer's account, send the reply. One-way tier.
That is three gated actions out of nine, and two of them can be demoted with work. The refund stays one-way because money moved. The account deletion becomes cheap the moment you soft-delete. The reply becomes cheap the moment you add a five-minute send delay.
After one afternoon of work, the agent has exactly one gate, on refunds. That is an agent people will actually leave running, because the gate is rare enough that they read it.
What this is not
This is not a substitute for scoping what the agent can reach at all. An agent with production database credentials it does not need is a problem that reversibility analysis will not solve, and sandboxing is the answer there. Reversibility tiers tell you where to put the human. Scope tells you what is in reach in the first place. You want both.
It is also not a claim that human review catches everything. Human in the loop degrades fast when the loop fires often, which is the entire argument for keeping the gated list short.
FAQ
What makes an AI agent action irreversible?
Four things: someone outside your system saw it, money moved, an external system committed to it, or the previous state was destroyed with no copy. Any one of those is enough.
Should I require approval for every agent action?
No. Frequent approvals train people to click through without reading. Gate only the actions you genuinely cannot undo, and convert the rest into reversible ones where you can.
How do I make a delete reversible?
Soft delete: mark the row deleted rather than removing it, and filter it out of reads. Hard deletion becomes a separate, scheduled job that runs after a retention window, which gives you a real recovery period.
What is the single highest-value change here?
A short outbound hold queue on anything the agent sends to a person. Sending is the most common one-way action in most systems and the easiest one to make reversible.
Does a spending cap replace gating purchases?
It limits the damage rather than preventing the action. Use both: the cap protects you when the gating logic has a gap, because it lives in a separate system that does not share the same bug.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


