Home / Tech / Anthropic Text Watermarking: Can Claude Secretly Mark AI-Generated Writing?

Anthropic Text Watermarking: Can Claude Secretly Mark AI-Generated Writing?

Anthropic Text Watermarking

Anthropic is adding an invisible watermark to text generated by future Claude models, creating a new way to estimate whether Claude was involved in producing or processing a piece of writing. The system does not insert hidden characters, add visible labels, or alter the text in an obvious way. Instead, it changes how Claude uses randomness when choosing between similarly suitable words.

Anthropic says the change is being introduced globally to help comply with the EU AI Act and related transparency requirements. The company also stresses an important limitation: a watermark does not prove that Claude wrote an entire piece of content. It only indicates that Claude was likely involved at some point.

What exactly is a text watermark?

Large language models generate text one token at a time. At many points, several possible words can produce essentially the same meaning.

For example, after:

“The weather today was cold and…”

the model might reasonably choose “overcast” or “grey.” Neither choice materially changes what the sentence means.

Normally, Claude can use a random process to decide between such possibilities. Anthropic’s watermarking system changes the source of that randomness.

Instead of relying on an arbitrary random number, Claude uses a secret key together with words appearing earlier in the text to determine which suitable option it selects. Across a long passage, those choices create a statistical pattern. The pattern is invisible to readers, but someone with the appropriate key can test whether the sequence of words is consistent with Claude’s watermarking process.

In simple terms:

Normal Claude output → sensible word choices + ordinary randomness

Watermarked Claude output → sensible word choices + key-controlled randomness

The important point is that Claude is not simply being forced to use particular words. It can still choose among words that make sense in context. The watermark changes how the random choice is made, not the fundamental meaning of the response.

Will readers be able to notice the watermark?

Anthropic says no.

The company states that watermarked and unwatermarked Claude responses should be indistinguishable to readers and that the technique has no practical impact on output quality or content. Internal testing reportedly found no measurable effect on creativity, readability or content quality. Anthropic also points to research behind the underlying technique, including testing by Google DeepMind that found no statistically significant difference in user ratings between watermarked and unwatermarked outputs.

That makes this fundamentally different from a visible watermark on a photograph or document.

There is:

  • no visible label;
  • no hidden Unicode character;
  • no additional text;
  • no extra tokens; and
  • no separate identifying information attached to the words themselves.

The watermark is essentially a statistical fingerprint rather than a hidden piece of text.

Which technology is Anthropic using?

Anthropic says Claude’s system is based on a version of SynthID-Text, a watermarking approach developed and published by Google DeepMind.

The method belongs to a broader family of watermarking techniques that use the model’s token-selection process to introduce a detectable statistical signal while preserving the quality of generated text.

That distinction matters because AI watermarking is often misunderstood as the model secretly inserting special symbols into its output.

That is not what is happening here.

The words themselves remain ordinary words. The signal emerges from the pattern of choices across many tokens.

What happens when Claude edits human writing?

This is where the system becomes much less straightforward.

The watermark applies to words Claude chooses.

If a person writes an article and asks Claude only to correct grammar and punctuation, most of the words remain the person’s original words. Claude therefore has relatively few opportunities to apply its watermarking mechanism.

As a result, the watermark may be too weak to detect reliably.

If Claude substantially rewrites the material, however, it makes many more word-selection decisions. The statistical signal can consequently become stronger.

This creates an important distinction:

Human text lightly proofread by Claude ≠ necessarily detectable as Claude-generated text

Human text substantially rewritten by Claude = more likely to produce a detectable Claude watermark

So a watermark cannot reliably answer the simplistic question, “Did AI write this?”

It answers a narrower question:

“How likely is it that Claude was involved in producing this text?”

What about factual writing?

Watermarking becomes weaker when there is little freedom in the model’s choice of words.

Consider a factual statement such as:

“Isaac Newton’s most famous work was called Principia…”

There is essentially one correct continuation: Mathematica.

Replacing it with another word would make the sentence wrong. There is therefore little room for the watermarking system to influence the choice without potentially damaging accuracy.

Anthropic says the same limitation applies to other passages where exact wording is necessary.

This means highly factual material can contain less watermark signal than ordinary prose.

That is an important weakness. A watermark is strongest when the model has many acceptable choices and weakest when accuracy requires a specific answer.

And what about programming code?

Code presents an even clearer example.

When Claude generates code, changing a token that is technically required can break the program. In those situations, there is no safe alternative for the watermark to exploit.

Anthropic therefore says code generally contains less watermarking than other forms of text.

There can still be opportunities to watermark arbitrary language inside code—for example, comments—because comments can sometimes be phrased in several equally valid ways. But Anthropic says this should have a negligible effect on the actual code produced.

Does the watermark make Claude slower or more expensive?

According to Anthropic, no.

The watermark has a negligible effect on model speed, and because it does not generate additional tokens, Anthropic says it does not increase the price of serving or using the model.

Can the watermark identify the person who used Claude?

No.

Anthropic says the watermark contains no information that can identify an individual user, organization or conversation. It applies to Claude’s output rather than functioning as a personal tracking mechanism.

This is a critical distinction.

A detector may be able to estimate:

“Claude was probably involved.”

It cannot use the watermark itself to conclude:

“This specific person generated it.”

Can someone remove the watermark by editing the text?

To some extent, yes.

Anthropic acknowledges that light editing may not completely remove the watermark. But a complete rewrite in which essentially every word is replaced can eliminate the statistical pattern.

At that point, however, the company notes that it becomes debatable whether the resulting text should still be described as AI-generated.

This exposes one of the central limitations of AI watermarking.

A watermark is evidence of provenance, not an indestructible fingerprint.

It can strengthen confidence that Claude was involved, but it cannot create an absolute historical record of everything that happened to a document afterward.

What does a positive watermark actually prove?

Less than many people might assume.

A detected watermark can indicate that Claude was likely involved with the content at some point. It cannot distinguish between:

  • Claude writing the content from scratch;
  • Claude heavily rewriting human material;
  • Claude editing existing material; or
  • another form of substantial Claude involvement.

Anthropic explicitly says the watermark cannot distinguish “Claude wrote this” from “Claude heavily edited this.”

That makes the wording around AI detection extremely important.

Calling something “AI-generated” based solely on a Claude watermark could therefore overstate what the technology actually establishes.

A more accurate interpretation is:

“There is evidence that Claude was likely involved in producing this text.”

What about translations?

Translations are a stronger case for watermark detection.

Anthropic says a translation produced by Claude carries a watermark because Claude is choosing essentially every word of the translated text.

That is very different from asking Claude to correct a few punctuation errors in a document written entirely by a human.

The more words Claude actually chooses, the more opportunities there are for the statistical watermark to accumulate.

How will people check whether Claude was involved?

Anthropic says it plans to offer a watermark detection API, although implementation details were still being worked out at the time of its announcement.

This could eventually give platforms, researchers and other organizations a way to test text for evidence of Claude involvement.

But it should not be confused with conventional AI-detection software.

Traditional AI detectors generally look for stylistic or statistical characteristics associated with machine-generated writing. Anthropic’s watermark detector works differently: it can test whether the text contains the statistical signature created using Claude’s secret key.

That distinction is crucial.

AI detector: “Does this writing look like AI-generated text?”

Claude watermark detector: “Does this text statistically match the watermark Claude used?”

The second question is narrower—but potentially much more directly tied to provenance.

Does the watermark change ownership or authorship?

No.

Anthropic says the watermark does not determine ownership, establish authorship or change a user’s rights under its terms. It simply indicates that Claude may have been involved in processing the content.

That distinction could become increasingly important as AI becomes a normal part of writing workflows.

Imagine a journalist who uses Claude to translate an interview, an editor who uses it to restructure paragraphs, or a developer who uses it to rewrite comments.

A watermark could show that Claude was involved without proving that the underlying work was authored by AI.

What happens to images and other files?

Anthropic is taking a different approach for supported files.

When Claude creates certain files such as PNG, JPG or SVG files, Anthropic says it will attach a C2PA content credential to the file’s metadata. C2PA is an open industry standard designed to record information about where digital content came from.

Unlike the text watermark, this is not a hidden alteration of the content itself.

The credential is a cryptographically signed piece of metadata indicating that Claude was involved in creating or processing the file. Anthropic says it does not contain identifying information about the user.

So Anthropic is effectively using two different provenance mechanisms:

Text → statistical watermark

Supported files → C2PA content credentials

Why is Anthropic doing this now?

The immediate reason is regulatory compliance.

Anthropic says it is implementing watermarking to comply with the EU AI Act. The company also says it and other major AI providers signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026, with the code calling for methods to mark AI-generated text.

Anthropic is applying the watermark globally at launch because it does not currently have a durable way to restrict the mechanism by region.

That means the change is bigger than a European compliance feature.

It potentially changes how Claude-generated content can be identified across the internet, regardless of where the user is located.

The biggest limitation: a watermark is not a truth machine

This is where the marketing around AI provenance needs to be treated carefully.

A watermark can provide useful evidence, but it does not reconstruct the history of a document.

A positive result can suggest:

Claude was probably involved.

It cannot establish:

Claude wrote the entire thing.

And a negative result does not necessarily establish:

No AI was involved.

A short passage may not contain enough choices for reliable detection. Factual writing may provide too little freedom for the watermark to accumulate. Light editing may leave too little Claude-generated material to detect. And extensive rewriting can eventually destroy the signal.

In other words, watermarking is best understood as a probabilistic provenance signal, not a universal AI lie detector.

What this means for the future of AI-generated content

Anthropic’s approach represents a significant shift in the AI transparency debate.

For years, the central question was:

“Can we tell whether this text was written by AI?”

Watermarking changes the question to something more precise:

“Can we determine whether a particular AI system was likely involved in producing this content?”

That is a much more useful question for provenance.

It also exposes a more complicated reality. AI assistance is no longer binary. A person might write 95% of an article and use Claude for grammar. Another person might provide a five-word instruction and have Claude generate everything. Both workflows involve AI, but they are obviously not equivalent.

The watermark cannot resolve that distinction by itself.

And that may be the most important takeaway.

Anthropic’s watermark can provide evidence of Claude involvement, but it cannot decide what that involvement means.

As AI becomes embedded in writing, translation, coding, editing and creative work, the real debate will move beyond human versus machine.

The harder question will be:

How much AI involvement matters—and who gets to decide when it should be disclosed?

Tagged: