By using a specific cryptographic key to determine the "randomness" of word selection, Anthropic can later analyze a passage to calculate the probability that it originated from Claude.
The implementation is primarily driven by the EU AI Act, which requires AI providers to mark generated content to increase transparency.
This move matters because it establishes a standardized way to track AI-generated text across the industry without increasing costs or latency—the time it takes for the model to respond.
Because the watermark is embedded in the statistical pattern of the words themselves rather than through hidden characters, it remains detectable even if the text undergoes light editing.
However, the system has technical limits in specialized contexts like computer programming or factual reporting.
In these cases, there is often only one correct way to write a line of code or a specific fact, leaving the model with no "low-stakes" choices to embed the watermark.
While the watermark can confirm if Claude was likely involved in a translation or a long essay, it is less effective for short samples or instances where the AI only performed minor proofreading on human-written text.
To support this new standard, Anthropic plans to release a detection tool that allows users to check for the presence of these marks.