Does Claude's watermark solve the AI transparency problem?
Anthropic now watermarks text generated or processed by Claude. Here's how the hidden pattern works, its limitations and why the approach has drawn criticism.
As of Aug. 2, all text made with Claude will bear a word-pattern-based watermark that is detectable by machines but invisible to human eyes.
Does it matter if the image you see or text you read was created by a human or generative AI? In some contexts, it matters a great deal. That is why the EU introduced Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems on Aug. 2, prompting Anthropic to announce the watermarking strategy.
Anthropic has committed to applying its watermark to output from supported models wherever Claude is offered and to add marking support to Claude models released before Anthropic introduced the watermark.
How does Claude’s text watermarking work?
Claude’s approach operates at the model level, meaning it will be present regardless of the specific product or surface overlaid on Claude.
Claude creates the watermark by replacing the LLM’s random token selection with a specific, checkable pattern set by a secret key. The possessor of the key can determine if the text matches such a pattern to flag what shows strong indications of having been composed by Claude.
When Claude generates text, it interlaces an imperceptible watermark directly into the text itself. According to Anthropic, the inclusion of the watermark does not change the meaning, readability or quality of Claude's response. It persists through copying and pasting, and some editing.
How Claude marks supported filetypes
Claude's approach to watermarking generated file types such as a .svg, .png, or .jpg differs from the text-based approach.
In these cases, it attaches signed provenance metadata that follows the Coalition for Content Provenance and Authenticity (C2PA) open standard to establish content provenance. A signed metadata label signals that Claude processed the file and provides a way to detect whether the file has been tampered with.
The text watermark is not a piece of metadata or additional characters included in the text, it's a pattern in the text itself.
How does it compare to other watermarking approaches?
Anthropic did not invent this approach to flagging AI-generated content. Google DeepMind introduced SynthID -- its version of a proprietary statistical watermark that lives in generated text -- in 2023. Both are statistical/model-level approaches in which a hidden signal influences token selection to establish a pattern discernible to the detector in possession of the key. They also both persist through copy-paste..
Claude's approach and SynthID also share the same weaknesses:
- Rely on longer texts to establish the pattern. This approach should work on an article or blog post but likely would not flag a short social media post.
- Requires word selection. Claude might not generate the pattern for text that doesn't involve any unique word selection because it is made up solely of code, numbers or even lists of dates or proper names.
- Significant editing can break them. Even though the pattern is not explicitly stripped out by a copy and paste, if the text gets altered significantly in editing, the pattern might fade away.
The key differences between SynthID and Claude's approach are a matter of model specificity and scope.
- Specificity. Google's SynthID detector is based on its own secret key which only applies to Google-generated watermarks. It cannot detect Claude's. The converse is true as well. Anthropic's key only works for Claude, not for any other model.
- Scope. SynthID works for a wide range of files, including images, audio and video. In contrast, for anything other than text, Anthropic has stated that it will rely on the C2PA open standard instead of the token sampling approach. There is no proprietary secret key involved in the C2PA approach. Instead, it attaches identifying metadata to the file.
How C2PA compares
Anyone can strip C2PA metadata out with minimal effort. But it offers two significant upsides over the token sampling approach:
- The open standard enables independent verification without access to a secret key.
- The metadata can survive extensive revision.
Yay or nay on Anthropic’s announcement?
While transparency is generally considered a good thing, the reaction to Claude's announcement was not universally positive. Anthropic has taken note of some responses and addressed them in speaking with journalists and directing people to its FAQ in an X post on Aug. 14.
A good portion of those 1.8K comments represent people firmly in the nay group. Even some who say AI detection is both not new and necessary are not attributing the best intentions to Anthropic's move.
Donn Felker, a software developer, pointed out on X that the watermark could help AI models recognize and avoid training on AI-generated content – a circumstance that can degrade and homogenize outputs.
Despite espousing the potential benefits of the approach, a follow-up response to a comment on his post comes off as more critical:
"Yeah the Anthropic anti-bad guy stance isn't something I really believe in anymore. Especially after dealing with their enterprise sale process. They're very good at getting every penny they can."
The connotations of a secret key
On Aug. 11, Bill Gurley wrote on X: "The word "watermark" comes from photos and is visible by all. This is only identifiable by Anthropic. Once again, they are judge, jury and prosecutor. Lots of reasons to not use their products."
Anthropic must have recognized its exclusive possession of the key as a valid concern. A representative told Business Insider that it intended to release an API "that will allow users and third parties to check text for Claude's watermark themselves." That would appear to be a step in the right direction for any business that claims to be acting in the name of transparency.
A lack of nuance in Claude's watermarking
Critics of AI detectors often point out that they can flag wholly human-created text as containing signals of AI-generation due to patterns that LLMs learned to emulate. As Andrea Saez explains in the blog "Putting Claude's watermarking to the test," one should not mistake signals for absolute proof of AI authorship. Due to the common use of particular phrases and repeated words, it is possible that a detector might pick up false positives, even from a text that is 100% human-generated.
Real life examples of this abound.
Theo Browne, a software engineer and entrepreneur who posts popular tech content under the handle t3.gg, mentioned that contemporary content creators often have their content flagged as AI because the LLMs were trained on their content.
The fact that Claude seems to apply the watermark to everything it touches exacerbates the problem. Someone who composes a text and merely uses Claude to proofread or translate it would receive the same label as someone who had Claude create the text from scratch.
Simon Smith, executive vice president of generative AI at Klick Health, addressed that on X on Aug. 16: "To clarify, I don’t think most people are upset by HOW Anthropic is watermarking. It’s that they’re stamping "Generated by Claude" on everything, not just fully AI-generated articles, but literally everything Claude helps with, which goes far beyond the EU AI Act’s requirements."
Smith's concern was echoed in the subhead of an Ars Technica article on the news, "Claude’s new Scarlet Letter watermark is invisible -- for now." The subhead notes: "The mark flags anything Claude processed, even human writing it only edited."
The Business Insider article that gave Anthropic the opportunity to respond to concerns, the unnamed company representative dismissed Smith's concern about light Claude use triggering the statement. "A watermark shows that Claude processed text, not necessarily that it wrote it, and that the mark can remain after Claude has proofread, translated, or summarized content."
One wonders if the Anthropic representative is being deliberately obtuse here. Smith's and Ars Technica's argument is that using one label for any text touched by Claude is misleading.
Claude goes way beyond the letter of the law
Even though Anthropic claims that its watermarking is necessary to remain in compliance with the EU AI Act, it goes much further than required. Content that the EU law allows to remain label-free would still get the mark.
Even wholly AI-generated text can be exempt from the labeling requirement, as is the case for works of fiction or even marketing copy.
Only text intended "to inform the public on matters of public interest" is required to show AI labels (and there's even a loophole for that if a human editor signs off on it).
Concerns about copyright, data privacy and quality questions
There are additional copyright concerns as well. For example, the UK-based software engineering coach John Crickett posted on X:
Code from Claude Code is now watermarked. I wonder where that leaves the copyright ownership? If the code [is] AI generated, copyright cannot be claimed. If it's watermarked as AI generated, how can the author claim it had enough human input to claim copyright? How could this affect the valuation of tech companies?
Steven Sinofsky wrote on X that he was worried about digital privacy: "The real issue is data retention and your right to private thoughts free of a digital trail."
Is more authentication better?
The question that arises from this is: Does erring on the side of more watermarking make it impossible to pass off AI-generated content as human-generated? The answer appears to be no.
Saez, Ars Technica and Browne all point out that it is possible to obliterate Claude's watermark by pasting the text into a rival generative AI and telling it to edit it. "I just cannot fathom any method where you can watermark text that isn't either incredibly expensive to detect or incredibly cheap to work around," Browne said in a YouTube video.
Accordingly, he predicts "that we'll get over this watermarking phase relatively quickly," Instead he suggests focusing on marking authenticated human content.
Ariella Brown is a technology journalist with experience covering AI, blockchain, IoT and cybersecurity.