Dashboard

How to Prompt AI to Redact a Document Properly

Remove all personal information produces a document that looks redacted and is not. The identifiers that survive are the ones that only identify in combination.

Steve Jefferson
Steve Jefferson
Developer Advocate
5 September 20261 min read

Asking a model to "remove all personal information from this document" produces a document that looks redacted and is not. Names and email addresses disappear. What survives is the combination that identifies the person anyway: the job title, the office, the deal size, the date of the incident, the internal project codename. Anyone with context reads straight through it.

Redaction is a different task from anonymisation, and prompting for it well means naming what you want removed rather than gesturing at a category.

Why the obvious prompt fails

Run "redact all PII from this document" against a real internal memo and watch what happens. The model does a competent job on the things that look like personal data in isolation: full names, emails, phone numbers, addresses, account numbers. It leaves everything that is only identifying in combination.

A memo with the names removed that still reads "the Regional Director for the Nordics raised this after the Q3 renewal call with our second-largest logistics customer" identifies two organisations and one person to anyone inside either company. The model did what you asked. You asked the wrong thing.

The second failure is quieter and worse: the model paraphrases while redacting. You get a document that reads cleanly, has the sensitive parts removed, and no longer says exactly what the original said. For a document going to a regulator, a court, or a customer, a fluent paraphrase is a serious problem, and it is invisible unless you diff the two.

The two-pass pattern

Separate finding from removing. One pass identifies, one pass redacts, and you get to inspect the list in between.

Pass one: enumerate, do not edit

text
Below is a document. Do not rewrite or redact it.

List every span of text that could identify a person or an
organisation, either directly or in combination with other
spans. For each one output:
  - the exact text, quoted verbatim
  - the category (see list below)
  - direct or indirect

Categories: person name, role or job title, organisation name,
location, date, monetary amount, internal identifier or code,
system or product name, distinctive event description.

Include anything that is identifying only when combined with
another item in this list. Over-include. I will filter.

The categories are the working part. Enumerating them forces attention onto indirect identifiers, which is exactly what the vague prompt misses. Getting the model to produce a candidate list before acting, rather than acting directly, is what makes the whole thing inspectable.

You now have a table you can read. This is the point of the pattern. A human decides what actually needs to go, which is a judgement call about audience and risk that no prompt should be making on its own.

Pass two: replace exactly, change nothing else

text
Below is a document, and a list of spans to redact.

Replace each listed span with a placeholder in this form:
[PERSON-1], [ORG-1], [DATE-1], and so on. Reuse the same
placeholder for the same underlying entity throughout, so the
document stays readable.

Change nothing else. Do not rewrite, reorder, summarise,
correct grammar, or improve phrasing. Every character outside
the listed spans must be identical to the input.

Output the redacted document only.

Consistent placeholders matter. Replacing every name with [REDACTED] destroys the document's meaning; replacing them with [PERSON-1] and [PERSON-2] keeps the relationships intact, which is usually the point of sharing the document at all.

The "change nothing else" instruction is doing real work here, and it is worth being emphatic about it. Models default to being helpful, and helpful includes tidying. Structured output helps too if you are doing this at volume, since getting JSON output makes pass one machine-readable and the verification step below far easier.

Verify with a diff, not with a read

This is the step people skip, and it is the only one that catches the paraphrase failure.

Take the original and the redacted output. Diff them. Every changed region should correspond to something on your list from pass one. Anything else is the model editing when it was told not to.

For a short document, do this by eye. For anything longer than a page, do it mechanically. A word-level diff on two text files takes seconds and turns a trust problem into a check.

If the diff shows unauthorised edits, do not argue with the model about it. Rerun pass two with the offending behaviour named explicitly in the prompt. Asking the model to check its own work is a useful complement here, and a poor substitute for the diff.

What this pattern does not solve

Three honest limits.

It is not a compliance control. A two-pass prompt is a drafting aid that makes a human reviewer faster. If the redaction has legal weight, a human signs off on the output, every time.

Numbers survive re-identification better than you think. A salary figure, a headcount, or a contract value can identify an organisation uniquely inside a small market. Categorise monetary amounts as identifying by default and let the reviewer decide, rather than the other way round.

The document went to the model. Redacting a document with an AI tool means the unredacted document was transmitted, processed, and possibly logged. That is the tradeoff, and it is a real one. If the document is sensitive enough to need redaction, it is worth knowing your provider's retention policy first, and the prompt-side habits in prompting without leaking sensitive data apply doubly here. Teams that have already had a confidential paste incident tend to reach this conclusion the hard way.

A shorter version for low-stakes documents

If you are redacting a meeting note before sharing it internally, the full two-pass process is overkill. One prompt does fine:

text
Redact this document. Remove person names, organisation names,
and exact dates, replacing each with a consistent placeholder
like [PERSON-1]. Also flag, in a list at the end, anything that
would still identify someone in combination. Do not rewrite
anything else.

The trailing flag request is the part that matters. It gives you the indirect identifiers as a list without committing the model to act on them, which for an internal document is usually the right balance. The same instinct applies whenever you are pulling structure out of prose, which is what extracting data from a document does from the other direction. The rest of the toolkit lives in prompt engineering.

The same grounding discipline applies to summarizing rather than redacting, how to prompt AI to summarize a PDF accurately covers the citation technique that catches invented detail.

FAQ

Can AI redaction replace a human reviewer?

No. It reduces the reviewer's workload substantially and does not remove the review. Treat the output as a first draft with a high hit rate, not as a finished document.

Why does the model keep rewriting sentences I did not ask it to change?

Because rewriting is its default mode and redaction is an editing task. The fix is an explicit instruction that every character outside the listed spans must be identical, plus a diff to confirm it happened.

How do I redact a scanned PDF or an image?

Extract the text first, redact the text, and regenerate the document. Never paint a black box over an image-based PDF and consider it done, since the underlying text layer often survives.

Should I use placeholders or just delete the text?

Placeholders, almost always. Deleting breaks the document's readability and makes it impossible to tell whether two mentions referred to the same person.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.