Anthropic has revealed how its controversial new change to Claude will actually work.
Last week, the AI company revealed that it would start leaving an invisible watermark in text, so that it would show up as being automatically written when fed through a machine that could spot that mark. But it did not say exactly how it would work – until now.
Anthropic has said that the feature is being introduced in response to new European rules, though it would roll it out across the world. It also suggested that it would help with “transparency”, by allowing people to understand where a text had come from.
But it led to intense criticism from some users, who said that it was unfair to punish people who rely on AI systems to write for them. Others argued that it would be easy to circumvent, could lead to accusations against people who had only used Claude to check or proofread their writing, and that it could undermine the quality of the system’s output.
Now, Anthropic has revealed how the system will work. And it stressed once again that the watermark “does not have any practical impact on the quality or content of Claude’s outputs”, that it would not be obvious to readers, as well as confronting some of those other criticisms.
In short, and as many had speculated, the system works by favouring some words over others, in a way that is not expected to be obvious to readers but will be statistically significant enough that it can be picked up by checking software.
It used the example of a sentence such as “The weather today was cold and…”. The next word might be “overcast” or “grey” – and while readers may not make much of a distinction between the two, and would be unaware even that the choice was made, the decision to favour one over the other could be one part of the watermark that checking software would be able to look for.
“Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses,” Anthropic wrote. “That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it.
“When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key.”
If there enough of those choices made, then software will be able to pick them up, and show that the text was made with Claude, Anthropic said.
The system is similar to Google’s SynthID, which Anthropic explicitly said had influenced its choice of the technique. That tool works in much the same way and was first introduced in a Nature article in 2024, though it was inspired by previous work.
Anthropic also addressed widespread fears that the watermarking could be used to accuse people of entirely creating text with AI when in fact they had only used it to edit their work. It said that the watermark “only applies to words Claude chooses”, and so if a piece of text had genuinely only been “lightly edited” then there would be too few words changed to actually leave behind the watermark.
“The more Claude writes, the more decisions it has to make, and the more space there is for a watermark,” it said.
But that new explanation of how the system works and Anthropic’s attempts to address criticism have only led to yet more outrage among some users. Tech commentator John Gruber for instance argued that the system would “adulterate and corrupt” Claude’s output.
“The idea that anything other than my needs should factor into the generation of text for me is patently offensive,” he argued in a long post on his blog. He argued that the feature would make the tool “worse” in an angry post that told Anthropic to “get f***ed”.
Anthropic also stressed that Claude is not the only AI system that will be introducing such watermarking, given that the EU regulations affect all AI providers. But there is nothing in those rules that require such watermarks to be present across the world, as Anthropic has chosen to do.
