I have written about AI slop and AI fakes as part of my work here on the ‘metaphor observatory’ in the era of generative AI. So, when I saw posts on Bluesky about ‘watermarks’, I pricked up my ears. The buzz was all about Anthropic adding ‘watermarks’ to Claude so that one can see what is written by humans and what is written by AI and to distinguish between what is real and what is fake.
Watermarks have been discussed in this context since 2021/22 mostly for images but also for text used to verify “the authenticity of the output or the identity or characteristics of its provenance, modifications, or conveyance”, as Zhang et al. pointed out in a 2023 preprint. For me the word ‘watermark’ still just conjures up a faint design or pattern pressed into paper or overlaid on digital media to show proof of human ownership or authorship. How does that work with AI I wondered where we want proof of AI authorship.
Claude doesn’t write on real paper. It’s all digital. So, we are in the realm of metaphor where some aspects of the concept of a physical watermark are mapped, one would think, onto the concept of a non-physical watermark in synthetic text. Let’s see how that plays out. It’s complicated.
Metaphors and mapping
Before we start, a quick reminder of what a metaphor does. It maps aspects of a cognitive or conceptual source domain onto a cognitive or conceptual target domain. Normally, the source domain is somewhat more familiar than the target one. Normally, this mapping process helps us to get some understanding of the second, the target domain, the novel phenomenon or novel aspect of a phenomenon.
Take the metaphor “Juliet is the sun”. Here we map familiar aspects of the sun (radiance, warmth etc.) onto a person and see the person in a new light. Or take “Genes are blueprints”. Here we map aspects of architectural drawings onto biological entities and come to some, now disputed, understanding of how genes work to make us who we are. In the following we’ll try to see what happens when we map ‘watermark’, the source domain onto what happens inside Claude, the target domain and what this means for at least my own understanding or not of what’s going on there.
I’ll start at the beginning with Anthropic’s announcement, then, after a quick detour into the past, I’ll try to explain what watermarks do in the world of AI (but I am really not an expert on that!); after that I’ll map the old and material onto the new and immaterial and see how it all matches up or not in terms of metaphorical mappings.
Anthropic’s announcement
On 11th August 2026 Anthropic announced that its Claude AI models would from then on embed invisible watermarks in text outputs and signed metadata in image files to comply with the European Union’s AI Act’s Article 50(2) Code of Practice on Transparency of AI-Generated Content, which took effect on August 2, 2026. It also announced that it was rolling the watermark out globally, not just for EU users.
This caused a stir on social media, including speculations about whether a change of one word or comma by Claude in something initially written by a human author would trigger the watermark. There was also talk about FUD, which I found out stands for Fear, Uncertainty and Doubt. Somebody even talked about this “generating angst among slopists”. I can sympathise with that, although I am not a ‘slopist’, but I use Claude occasionally.
As a non-AI-tech-expert, I began to wonder what all this watermark talk was about and as a historian of language and science I went back to the roots.
Watermarks in history
Adding a watermark to papers or stamps or banknotes is an old technology. It is a “physical design pressed into wet paper during manufacturing to prove authenticity”.
According to Wikipedia, watermarking originated in 1282 in Fabriano, Italy, where papermakers sewed wire shapes onto the mould used to form each sheet. The wire displaced pulp fibres at the point of contact, leaving that patch of paper thinner and so, once dry, faintly translucent. Held to the light, the mark appears, but in ordinary use, it’s invisible.
As somebody explained in a letter from 1787 (Oxford English Dictionary): “The paper-marks are those figures formed by wires, on the sieve at the bottom of the mould in which the paper is made and are impressed on it in its pulpy state…”. Watermarks were used to discourage counterfeiting, ensure copyright security and demonstrate provenance.
Watermarks in Claude
Watermarks are used in Claude to mark AI-generated text and to distinguish it from human-generated text. But how are these ‘watermarks’ ‘made’? Not with a mould, and a sieve, and wires for sure. It’s rather complex. It’s actually “pure probabilistic maths”, as Satyam Sahu explains on Medium. And that’s where the metaphorical mapping gets complicated; for me at least.
Let’s start with a simple definition provided by Zhang et al.: “A watermarking scheme consists of a generation algorithm that is a modified version of the model in which the signal is planted and a detection algorithm that can detect whether a piece of output came from the watermarked model.” I then asked Claude to describe what’s going on there in simple terms. Here goes (see also here and here for background, but there is so much more) (and I bet this will earn me a watermark)…
When Claude writes a sentence, at every single word it isn’t just picking “the best word”. Instead, it has got a shortlist of several good candidates and picks one, with a bit of built-in randomness. That randomness has always been there; it’s part of why the same prompt can get you slightly different phrasing each time.
The watermark doesn’t add anything new to the text. It just rigs which random choice gets made. Instead of that word-by-word coin-flip being genuinely random, it’s steered by a hidden key and the pattern of which words got picked, across a whole essay, encodes a signature. The pattern is undetectable to the reader, but detectable to anyone who has a key that encodes it. Nothing is visibly different. No word looks ‘wrong’. It is the same kind of writing Claude would produce anyway, just with the dice loaded in a traceable way.
What the watermark or SynthID-Text does is closer to steganography than to watermarking: it’s not adding a mark, it’s hiding a statistical signal inside a choice that must be made anyway (which word comes next).
So now we know how paper watermarking worked in the past and we know how the statistical text watermarking works in Claude (sort of!). We know something about the source domain, and we know something about the target domain. The physical/paper source domain is mapped onto statistical token-selection target domain. But what about that mapping?
Mapping across domains
What we got here is a rather complex metaphor, where some mappings work, some don’t and where the overall mapping might be rather difficult to recreate for ordinary users of AIs.
In the old technology the mark is made from the same substance as the paper, not added to it. In the new technology, the mark is made through token-probability shaping, not appended metadata – so there is a match here. In the old technology the mark is literally visible when held to the light; it’s invisible in normal use but revealed under specific viewing conditions. In the new technology the mark is imperceptible to the reader, but detectable by anyone who has a key – again we have a match here. Overall then there is source-target domain mapping between light and cryptographic key – it’s all about ‘transparency’ after all.
What puzzles me is how this metaphor can work for ordinary users who have no clue about steganography, statistical hypothesis testing and so on. Normally, in a metaphor that works well you map knowledge from one domain onto another and you succeed because you know or think you know a little bit about both. But here, most people just stay on one side of the mapping with the knowledge of physical watermarks, but they can’t make the jump into the target domain as the knowledge is too deep.
However, the watermark metaphor still makes something alien (keyed pseudorandom sampling bias, verified by statistical significance testing) feel familiar, physical and graspable. But… the source domain (ink, paper, light) supplies a complete, satisfying mental model that has nothing structurally in common with the target domain except the vague idea “something hidden that proves origin”. There are also some more glaring mismatches.
Mismatches in mapping
The Fabriano watermarks from the 13th century were proud identifying marks; they were the marks of craftsmen and of a mill, a mark of quality and origin. The AI watermark is more of a compelled disclosure, and I guess nobody would be proud to get one. It is a mark of misbehaviour, sloppiness, fraud and fakeness, rather than a trademark or a sign of craftsmanship.
And there is another mismatch. With the old technology a paper mill produced unique sheets of paper. With AI text we are dealing quite often with collaboratively produced output, an intricate blend of Claude writing and human writing. That makes the watermarking much more complex.
Not every text produced in collaboration with Claude is the same. Overlaps between human writing and AI writing can range from AI writing the whole thing, to AI just adding a comma, with a totally blended text in between that emerges from many rounds of writing and editing. How does the watermarking process work here?
As far as I can make out from reading, for example, Bruce Gil’s article in Gizmodo (see also article by Marc Bara here), it works best on longer, more open-ended writing, such as an essay, a script or an email, and it performs less well on narrowly factual answers, because there is less room to nudge word choice without risking accuracy. Its confidence drops if the text is heavily rewritten or translated into another language. This means that Claude-generated text may “lose the signal if it is heavily edited, paraphrased, translated, or mixed with other writing” (Gil) and that means that the watermark becomes, it seems, fragile exactly at the point where human co-authors assert their authority.
A placebo metaphor?
As Satyam Sahu said in his post about Anthropic’s announcement: “When most of us hear the word ‘watermark’, we picture a faint logo stamped across a photograph, micro-steganography hidden in file headers, or a subtle shift in pixel values.” To use this same metaphor for statistical text watermarking in all its complexity, can feel rather confusing. So what about the ‘watermark’ metaphor? Does it work? Probably not. It sort of breaks down when you look too closely at it.
One can almost call this a ‘placebo metaphor’. Rather than conveying or increasing knowledge of the complex target domain, it provides an illusion of knowledge. Anthropic’s own language leans into this illusion when they say, using a papermaking verb: “When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.”
Is the watermark metaphor a case of what Travis LaCroix et al. called strategic polysemy or ‘glosslighting’ by which they mean using “familiar words in a technical sense in a way that still evokes their everyday meaning, while keeping the option to fall back on the narrower definition when challenged”. Perhaps not. Should one use a more honest phrasing, such as ‘invisible signature’, ‘steganographic signature’, ‘machine readable text mark’, ‘machine detectable pattern’ rather than a metaphor that vanishes during use? Hmmmm…
Metaphors and their limits
Metaphors are a technology for thinking. They help us get to grips with something unknown through the lens of the known. But this technology has its limit. Let’s look at some examples.
Some metaphors successfully, although never completely, illuminate a topic (target domain). For example, the metaphor of ‘the big bang’ enables us to visualise the beginning of the universe. Some metaphors enable us to suddenly see a phenomenon in a new light and lead to productive discoveries, as in seeing the gene as a code. For some metaphors the light fades over time as the target domain is increasingly better understood, as in seeing the gene as a blueprint. In some cases, scientific phenomena are not sufficiently understood yet and we are groping for metaphors in the dark. In all cases of metaphorical mapping there must be at least some knowledge of both source and target domain for the mapping not to fall flat. There is always a tussle between the limits of our knowledge and the limits of our language.
What about ‘watermark’, a metaphor that I picked up with my metaphorical radar ears for the AI metaphor observatory? Here we have a scientific or technological phenomenon that is completely understood inside out but the metaphor that is supposed to make it familiar to those who are not familiar with it only scratches the surface. It means something for those with deep technical knowledge of probabilistic maths but confuses those of us who don’t possess that knowledge. Is it a good metaphor, a bad metaphor or a placebo metaphor?
Image: Wikimedia Commons: The image shows a piece of antique laid paper backlit to reveal an ornate coat-of-arms or heraldic watermark, showcasing the translucent chain and laid lines typical of historical handmade papermaking.

Leave a Reply