Anthropic has announced that upcoming Claude models will embed a watermark in the text they generate, and published an FAQ covering both the mechanism and its limits[1]. Readers cannot tell the difference, and the method adds no hidden characters and no extra tokens. The change responds to EU transparency rules that took effect on August 2, and Anthropic says it is applying the watermark worldwide[1][2].

Swapping the Source of Randomness

A large language model picks one word at a time. After "The weather today was cold and," the next word is very unlikely to be "sugary," but "overcast" and "grey" are both plausible and the sentence means much the same either way. In cases like that, a random number settles the choice[1].

Watermarking works on exactly those low-stakes choices. Instead of an arbitrary random number generator, the model uses a key plus a few preceding words to decide which candidate to take. The words are still random, but anyone holding the key can compare the resulting sequence and estimate how likely it is that Claude was involved[1].

Anthropic stresses that the method does not push Claude toward words it would never have considered, and that nothing is embedded in the text and no invisible characters are inserted[1].

The approach is a version of SynthID-Text, published by Google DeepMind in Nature in 2024, and belongs to a family of techniques tracing back to a 2022 proposal by Scott Aaronson[1][3].

No Effect on Quality, Speed, or Price

Anthropic says internal testing found no impact on the content, creativity, or readability of Claude's output[1]. The SynthID-Text paper reported the same: Google DeepMind served a watermarked model to a portion of Gemini traffic and compared thumbs-up and thumbs-down ratings, finding no statistically significant difference. The analysis covered roughly 20 million responses[1][3].

Because the method produces no extra tokens, serving costs and prices are unchanged, and the effect on speed is negligible[1].

Anthropic offers a Monopoly analogy. If players replaced dice with successive digits of pi, the experience and the outcome of the game would be no different. But reviewing every move afterwards would reveal that pi had been used. A text watermark sits in the same relationship to the text[1].

The Limits Are Spelled Out

What stands out in the FAQ is how concretely the limitations are described.

The key can only answer how likely it is that Claude partly wrote a passage. It does not prove that text was human-written, and it cannot identify text produced by a different AI, which would use a different key or an entirely different scheme[1].

Short samples are difficult. Fewer word choices mean less information, so confidence rises as a passage gets longer[1].

Factual writing carries a sparser watermark. After "Isaac Newton's most famous work was called Principia," only "Mathematica" is correct, leaving the watermark nothing to act on. Proofreading is similar: if Claude is asked to fix only grammar and punctuation, the watermark can live only in that handful of corrections[1].

Code is another gap. Where an exact output is required, the nudge is not applied, for the same reason that "2 + 2 =" has only one sensible continuation. Comments inside code can carry a watermark because word choice is freer there, but the effect on the code itself is negligible[1].

On removal, Anthropic concedes that light editing probably will not erase the watermark while a full rewrite replacing every word will, adding that at that point it is arguable whether the text is still AI-generated[1]. Translations do carry a watermark, because Claude chooses every word[1].

It Cannot Be Traced to a Person

The watermark applies to Claude and its outputs and carries no information identifying a user. Neither the watermark nor the key allows anyone to recover details about a person, an organization, or their chats[1]. Ownership of an output and legal responsibility for it are unchanged[1].

TechCrunch, citing Anthropic's support page, notes that because the watermark is part of the text it travels through copy and paste and may persist through some editing, and that it is applied at the model level, so it appears in text coming from the Claude platform API, Claude, Claude Code, Claude Cowork, and Claude Tag alike[2].

Files are handled differently. For supported formats such as .png, .jpg, and .svg, Claude attaches a cryptographically signed note in the file's metadata using the C2PA open standard. Nothing inside the file changes, which makes this a different mechanism from a watermark[1].

The approach also differs from third-party AI detection services such as Pangram. Without the key, those tools rely on the stylistic tells of AI phrasing. Anthropic cites the construction "this isn't [X], it's [Y]" and heavy use of the word "quietly" as examples[1].

EU Rules Started It, and Older Models Are Next

The trigger is the EU AI Act. In July 2026, Anthropic joined roughly 190 signatories to the EU Code of Practice on Transparency of AI-Generated Content[1]. The transparency code took effect on August 2 and requires providers serving the EU market to mark AI-generated content in a way other systems can identify[2].

Anthropic applied the watermark globally at launch because it does not yet have a durable way to scope it by region[1]. Models released before August 2 fall under a transition period, and Anthropic says it will extend coverage to them over the coming months[1]. A watermark detection API is also in preparation[1].

Black Forest Labs, Google, Meta, Microsoft, OpenAI, and Synthesia have committed to the same code[2]. This is not a Claude-only effort but part of a broader move in which each provider marks its own outputs[1][2].

Summary

Anthropic will watermark text generated by Claude using a SynthID-Text style method that changes only the source of randomness behind word choice, and says the change is invisible to readers and neutral for quality, speed, and price. The company also documents where the method falls short: short passages, factual statements, code, and proofreading, and the fact that detection only indicates possible involvement. It cannot identify individuals, and coverage for older models plus a detection API remain outstanding. How a marking scheme born of regulation reshapes the handling of AI-generated content is still an open question.

[1] Source: https://www.anthropic.com/news/claude-text-watermark

[2] Source: https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/

[3] Source: https://www.nature.com/articles/s41586-024-08025-4