Can AI Text Watermarks Be Removed?
Anthropic will watermark future Claude output using a version of SynthID-Text. Here is how the mechanism works, why the EU pushed it, and whether it can actually be stripped out.
Can AI Text Watermarks Be Removed?
Can AI text watermarks be removed? Not reliably, and not without giving up most of the fluency gains that made AI writing worth using in the first place. On August 14, 2026, Anthropic announced that future Claude models will embed an invisible watermark in generated text, describing it as a version of SynthID-Text, the technique Google DeepMind published in a 2024 Nature paper. The watermark adds no extra characters, no hidden tokens, and no visible artifacts. It works by steering ordinary word choices into a statistical pattern, one that can survive light editing but breaks down under heavy rewriting.
What Anthropic Announced on August 14
Anthropic's announcement says future Claude models will carry the watermark by default, with no change to output quality or generation speed. It is not retroactive across the board. Claude models launched before August 2, 2026 will get watermarking rolled out "over the coming months," not immediately, so plenty of existing Claude output carries no watermark at all. Anthropic also said it plans to ship a detection API so third parties, not just Anthropic, can check whether a passage carries the pattern.
The timing lines up with regulation, not a sudden burst of goodwill. Anthropic signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026, one of roughly 190 signatories to a code that asks AI providers to mark the content their systems produce. This is AI text watermarking arriving because of a compliance deadline, not a research team deciding it would be nice to have. Full context on what that code requires and when is in our EU AI Act transparency rules explainer.
How the Claude Text Watermark Actually Works
This is the same family of technique as SynthID-Text, and it is worth understanding the mechanism because it explains both why the watermark is hard to spot and why it is not bulletproof.
During generation, a model constantly faces small decisions where several words would work equally well, "quick" versus "fast," "start" versus "begin." A cryptographic key, combined with the words that came immediately before, quietly biases which of those equally good options gets picked. No single sentence looks unusual. Over enough words, though, the pattern of choices becomes statistically detectable to anyone holding the key, without changing meaning, tone, or accuracy in any way a reader would notice.
This is not unique to text. OpenAI has taken a related approach with SynthID watermarking for AI-generated voice audio, marking a different channel with a similar idea: encode a detectable signal into the output itself rather than a visible tag.
Can the Watermark Be Stripped Out?
Here is the honest answer, not the reassuring one: paraphrasing, translating a passage and translating it back, or heavily rewriting it by hand can break the statistical signal, because those actions replace enough of the original word choices that the pattern washes out. No watermarking scheme any lab has announced, including this one, claims to survive a determined rewrite.
But that is a hollow win. If you paraphrase or retype AI output aggressively enough to strip a word-choice watermark, you have also stripped away most of the reason to use AI writing in the first place, the speed and the fluency. You are back to editing prose sentence by sentence, which is close to writing it yourself. There is no method that removes the watermark while keeping the output untouched, because the watermark and the output are the same words.
It is also worth knowing what the watermark cannot do, even when it survives. Anthropic has been explicit that it cannot distinguish "Claude wrote this" from "Claude heavily edited this." It performs poorly on short samples, where there are not enough word choices to build a reliable pattern, and on dense factual text, where there is little room to vary phrasing anyway. A two-sentence email or a table of specifications will not carry a strong signal either way.
What This Means If You Write With AI
If you use Claude, or any model, to draft client work, blog posts, or marketing copy, the practical move is not evasion. It is disclosure that matches whatever your client, platform, or jurisdiction actually requires. Watermark detection is going to get easier over the next year, not harder, as detection APIs ship and more providers sign on to codes like the EU one. Treating a watermark as a scoreboard to beat is a losing long-term bet, treating it as a reason to be upfront about how you work is not.
Two limits are worth keeping in mind day to day. First, watermarking says nothing about whether the content is accurate, only whether it was likely generated by a specific model. Those are separate questions, and if you are worried about the second one, our guide on how to tell if an AI answer is hallucinated covers it directly. Second, this is a fast-moving area: model providers change watermarking policy, detection tooling, and disclosure rules often enough that it is worth checking in periodically rather than assuming today's rules hold in six months, which is exactly the kind of thing our guide to keeping up with AI news is built for.
FAQ
Does ChatGPT or Gemini also watermark text?
Google DeepMind, which originally published SynthID-Text, has applied the technique to Gemini output since the method became public in 2024. OpenAI has watermarked AI-generated voice audio using its own SynthID-based approach, but has not announced an equivalent text watermark for ChatGPT as of this writing.
How can I detect AI-generated text?
Not reliably on your own. General-purpose AI detection tools have a documented history of false positives and false negatives, especially on edited text. Anthropic's planned detection API will only catch Claude output specifically, only once a given model has watermarking enabled, and only on text that has not been rewritten heavily.
Does watermarking affect the quality of Claude's output?
No. Anthropic's design only nudges between word choices that were already equally valid, so it does not change meaning, tone, accuracy, or generation speed. That is the entire point of steering-based watermarking over cruder approaches, like appending a marker or altering formatting.
What happens to older Claude conversations and models?
Nothing retroactively. Text already generated before a given model gets watermarking enabled was never marked and will not be marked after the fact. Claude models that launched before August 2, 2026 are getting the feature over the following months, per Anthropic, not on a single fixed date.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


