透

Building the invisible watermark
hidden in LLM text,
with my own hands

Claude has started watermarking its generated text. How does that actually work β€” and does it really leave quality untouched while surviving edits? I implemented the token-space schemes (KGW, Gumbel-Max) and an embedding-space scheme (SemStamp-style) from scratch, and verified generation, detection, and attacks entirely with local models on a Mac. All sample texts below are unedited model outputs.

TL;DR

In three lines

The watermark lives in the interchangeable choices an LLM makes when picking the next token β€” places where either word would be natural. Each single choice reveals nothing, but summed over hundreds of tokens the key-driven bias produces a statistical signature (a z-score) that human text cannot exhibit. Token-level watermarks reliably die after one paraphrase; the semantic-space watermark degrades but can survive β€” and its survival hinges on how paraphrase-stable the sentence embedder is.

z = 4.02
Detection score of the SemStamp-style watermarked essay (Japanese setup, paper-faithful chained regions). The unwatermarked control passes at z = 1.83
z = 4.08
Score after paraphrasing that essay with an unwatermarked LLM. Across repeated runs the post-paraphrase score ranged z = -0.8 to 4.4 β€” guaranteed survival needs the paper's robust encoder
14.24 β†’ -0.92
The token-space watermark (KGW, English) is wiped out by a single paraphrase pass β€” reliably, in every trial
01 β€” Results

Results: watermarked and unwatermarked, side by side

Example 1 β€” Semantic watermark (SemStamp-style): degrades, but does not collapse

To make the pair directly comparable, the prompt pins the structure: both essays must open with the exact sentence "Why do people enjoy walking in the rain?" and cover, in order, the sound of rain, the smell and the air, and the calming effect on the mind. Generated in English by Qwen2.5-0.5B-instruct with all-MiniLM-L6-v2 sentence embeddings (fixed-region variant; see the fidelity section for the chained variant). Within that shared skeleton, the wording differs because sampling differs on every run β€” not because of the watermark. The watermark leaves no visible trace; it exists only in the statistic a key-holding detector can compute. A 0.5B model writes rough English; the texts are shown unedited.

No watermarkz = -1.33 / not detected

"Why do people enjoy walking in the rain?" is a question that speaks to our senses, emotions, and experiences. For deciding when to go for a walk during rainy weather is a personal decision based on various factors such as the climate, personal preferences, or even the weather forecast. Firstly, there's the auditory aspect of rainfall, which is often more prevalent than visual. The rhythmic thumping of waterfalls creates a soothing background noise that can help in calming one’s mind, making it easier to relax during a rainy day. Raindrops falling on wet surfaces create a sense of moisture and humidity, which can contribute to feelings of calmness and tranquility. Secondly, the tactile experience of walking in the rain adds an extra layer of sensory stimulation. The ground underfoot feels like carpeted grass beneath your feet, adding a touch of serenity to the already pleasant texture of wet earth. This tactile engagement also helps to increase body awareness and connection with nature around you. Lastly, the mental benefits associated with enjoying outdoor activities are equally significant. The feeling of being immersed in fresh air, surrounded by lush greenery, and feeling lighter overall during rainy days leaves many people feeling refreshed and rejuvenated.

12 sentences. The fraction of sentences landing in the key's "valid regions" stays at the chance level (expected 25%)
SemStamp-watermarkedz = 4.96 / detected

Why Do People Enjoy Walking in the Rain? People often seek out activities that offer fresh, natural environments to unwind amidst the hustle and bustle of urban life. Among these serene pursuits is one that has garnered widespread appealβ€”walking through the rain. The allure of nature's embrace of the elements, combined with its ability to bring a refreshing freshness, draws countless individuals into this outdoor activity. One of the primary reasons why walking in the rain evokes such fond memories is its captivating combination of sound and sensory experience. The rhythmic tapping of water as it falls against your feet can be a soothing lullaby, echoing across the landscape and providing a gentle yet constant melody that enhances the tranquility of the moment. This auditory cue creates an atmosphere of peace and calmness, allowing for introspection and a sense of well-being. Moreover, the smell of wet earth is another delightful aspect of walking in the rain. As moisture seeps through the soles of your shoes, the scent of damp grass and leaves fills the air, making you feel refreshed and rejuvenated. This pleasant aroma complements the sounds around you, creating a harmonious ambiance that encourages a relaxed state of mind. The air itself becomes particularly comforting during rainy days. The misty drizzle blankets the landscape, painting everything in a mellow hue of verdant greens and soft pinks. These natural elements not only enhance the aesthetic appeal but also provide a soothing contrast to the harsh exterior of modern cities.

13 sentences. Rejection-sampling tries per sentence: [17, 7, 1, 2, 11, 2, 9, 2, 1, 2, 17, 1, 15] (17 = fallback after hitting the retry cap). A quality filter rejects candidates containing markup or non-English fragments

I then had an unwatermarked model (the same LLM with watermarking switched off) rewrite the watermarked essay sentence by sentence, keeping each sentence's meaning β€” simulating an attacker laundering the text through a clean LLM, which adds no watermark of its own. The wording changed substantially, and β€”

↑ after paraphrasingz = 4.01 / still detected

Why do folks tend to enjoy wandering down below while it rains? People often enjoy pursuing activities that place a strong emphasis on finding environments free from modern noise and chaos, such as forests, vineyards, or even waking up early and venturing out into the evening fields. Among these idyllic activities is one that has gained widespread loveβ€”walking amidst the drizzle. The allure of nature's embrace of the elements and its capacity to refresh one quickly becomes irresistible, attracting and inspiring countless individuals to pursue this outdoor activity. A primary reason why walking in the rain evokes cherished memories is its striking blend of auditory and sensory delights. The rhythmic tapping of water as it falls against your feet can be a soothing lullaby. It echoes across the landscape, and it provides a gentle yet constant melody that enhances the moment's tranquility. This auditory stimulation produces a tranquil ambiance, promoting mindfulness and tranquility. In addition, the fragrant odors of damp soil are another delightful component of exploring in rain. When moisture seeps through the soles of your shoes, a pleasant aroma of damp grass and leaves fills the air, making you feel revitalized and refreshed. This pleasant aroma complements the sounds around you, elevating the mood and inducing a serene state of mind. On rainy days, the air feels particularly comforting. The drizzle-covered landscape descends in a palette of verdant greens and soft pinks. These natural elements not only grace the aesthetic decor but also serve as a natural counterpoint to the harsh exterior urban landscapes.

This run survived (z = 4.01) β€” but with an off-the-shelf English embedder (21% signature-bit flips under paraphrase) survival is probabilistic: earlier English runs degraded to z = 1.12 and -0.82. The Japanese setup with the more stable ruri embedder behaves the same way at a higher level β€” see the fidelity section

Example 2 β€” Token-space watermark (KGW): detected strongly, erased reliably

The same English prompt watermarked with KGW (a logit bonus on a key-derived "green list" of tokens, Ξ΄ = 2). The detector calls it with overwhelming confidence β€” and the same paraphrase attack erases it completely, in every trial. Note the prose itself is noticeably degraded ("cuisinarily encode", "the glands of your glands"): on a small model, the logit bonus visibly distorts generation, which is exactly the quality cost the distortion-free schemes were invented to avoid.

KGW-watermarkedz = 14.24 / detected

In damp, warm weather, there is no glory of walking like the soft embrace of the rain. It is the infinitude of the embrace that is unparalleled in nature. Raindrops falling from astronomical freedom are also serene. Rain is a mystery, sublime, mysterious.

Once in the headlands and orienteering coordinates of a rain forest or a quad of rain forests entwine cuisinarily encode the senses of the individual as if existing out of and from the mediocre fusion of public consciousness. Rain is intimately nurtured by our natural existence. Now, would spell β€œRain’s the only joy.” It’s the interpersonal nurture that we bloom with saturated in tenacity, that feels unparalleled in the social jungle where one walks too hard in a coffee-trough under a rain in the beauty of the substance.

Rain, the savor is without a doubt non-slip. The very brightness of the sun as it sheds a gentle and incomprehending embrace across them is accompanied by the powerful mix of fragrances and chemicals filled into an amber-stained sweat of your skin that spreads from the glands of your glands into your body. (…)

Paraphrasing this essay with the unwatermarked model yields z = -0.92 / destroyed β€” the watermark's substance is a bias in surface token choices, so a wholesale token replacement erases it

Example 3 β€” Where does the watermark live? Coloring every token

If two texts look indistinguishable, where is the watermark? At temperature 0 (deterministic decoding, no sampling noise) both versions even begin with the same eight tokens β€” "Walking in the rain can be a delightful" β€” before the logit bonus flips one close call and the texts fork. Running the KGW detector's own computation and coloring each token by whether the key favors it ("green") makes the watermark visible: the unwatermarked text is 20.0% green, near the chance level, while the watermarked one is 55.0% (Qwen2.5-0.5B-instruct, same prompt, temperature 0).

green-list token (favored by the key)red-list token
No watermark β€” 20.0% greenz = -1.63
Walking in the rain can be a delightful and relaxing experience for many people. The act of walking in the rain can be a refreshing change from the usual dry, sunny weather, and it can provide a sense of freedom and adventure. Walking in the rain can also be a way to connect with nature and appreciate the beauty of the world around us.

One of the reasons people enjoy walking in the rain is that it can be a great way to explore new places and discover new things. The rain can create a sense of mystery and wonder, and it can be a great opportunity to take in the sights and sounds of the world around you. Walking in the rain can also be a way to connect with nature and appreciate the beauty of the world around us.

Another reason people enjoy walking in the rain is that it can be a way to relax and unwind. The rain can create a sense of calm and tranquility, and it can be a great way to take a break from the stresses of everyday life.
KGW-watermarked β€” 55.0% greenz = 9.80
Walking in the rain can be a delightful experience for many people, especially when the weather is warm and the rain is light and gentle. The warm, wet weather can create a unique sense of serenity and comfort that is particularly appealing. Additionally, the sound of the rain can create a soothing background noise that helps to calm the mind and reduce stress.

Walking in the rain can also be a refreshing experience, especially if it is a warm day with a light breeze. The raindrops falling onto the grass and pavement can provide a refreshing contrast to the warm, humid air, and the sound of the rain can create a gentle rhythm that is soothing to the ears.

In addition to its calming effects, walking in the rain can also be a great way to explore the outdoors and get some fresh air. The gentle breeze and the sound of the rain can create a peaceful and invigorating atmosphere, making it a perfect activity for a relaxing and enjoyable walk.

Overall, people who enjoy walking in the rain find

No individual green token is special β€” "rain" and "walking" can be green or red depending on the key and the preceding four tokens at that position. The difference exists only in the aggregate over hundreds of tokens.

Example 4 β€” Quantitative experiments (both token-space schemes)

Detection z-scores with Qwen2.5-0.5B-instruct, averaged over 4 prompts Γ— 200 tokens. Only watermarked text is detected; unwatermarked output, human text (the opening of Austen's Pride and Prejudice), and detection with a wrong key all pass.

text under testKGW detector zGumbel detector z
no watermark (vanilla)-0.760.73
KGW-watermarked9.78-0.07
Gumbel-Max-watermarked-0.9034.99
wrong key-0.420.16
human text (Austen)-1.510.18

Attack experiments: random token substitution degrades the signal gradually, and a single paraphrase pass with an unwatermarked LLM destroys both schemes.

attackKGW zGumbel z
none5.5538.79
substitution 10%2.4519.61
substitution 30%-0.827.27
substitution 50%1.310.40
paraphrase (unwatermarked LLM)-1.45-0.66

On quality: KGW worsened mean NLL under the same model from 2.25 to 3.02 nats/token (a distorted distribution), while Gumbel-Max stayed at 2.77 β€” no degradation, exactly as the distortion-free construction promises.

02 β€” Mechanism

How it works

An LLM samples its next token from a probability distribution, and a passage offers many moments where "either this or that reads naturally". A watermark nudges these interchangeable choices with a rule derived from a secret key. One choice reveals nothing; hundreds of them, aggregated, reveal a statistically impossible habit. Detection needs no model access β€” only the key and the text.

TOKEN SPACE / 2023

KGW

Hash the preceding tokens with the key to split the vocabulary into a green list (25%) and a red list, then add +Ξ΄ to green logits. Detection is a z-test on the green-token rate. Simple, but it distorts the distribution and can hurt quality.

early implementation β†’ preserved in git history
TOKEN SPACE / DISTORTION-FREE

Gumbel-Max

Replace the sampling randomness with key-derived randoms r and pick argmax ri1/pi. The output distribution is provably unchanged β€” zero quality cost. Detection scores Ξ£ βˆ’log(1βˆ’r). SynthID-Text belongs to this family.

early implementation β†’ preserved in git history
EMBEDDING SPACE / 2024–

SemStamp-style (semantic)

Generate sentence by sentence, accepting only sentences whose embedding lands in "valid regions" of a key-derived LSH partition, with the valid set chained from the previous sentence's signature. Paraphrasing barely moves a sentence's meaning vector, so the watermark can survive.

what sukashi is now
generated sentence tends to stay in the region even after paraphrase valid region (key + previous sentence) the sentence-embedding space, partitioned by key-derived LSH hyperplanes generation loop 1. sample a candidate sentence 2. embed it 3. accept if in a valid region, otherwise go back to 1 (rejection sampling)
SemStamp-style semantic watermarking. Even when paraphrasing replaces the surface token sequence wholesale, the sentence's meaning vector stays nearby β€” as long as the embedder is paraphrase-stable.

Every detector is a hypothesis test

All three schemes reduce detection to one question: how many Οƒ does this text's statistic deviate from the no-watermark null? KGW and SemStamp use the normal approximation of a binomial test; Gumbel uses a sum of exponentials (Gamma null).

# KGW / SemStamp: green (or valid-region) rate test
z = (hits βˆ’ Ξ³Β·n) / √(nΒ·Ξ³Β·(1βˆ’Ξ³))

# Gumbel-Max: per-token score is Exp(1) under the null
S = Ξ£ βˆ’log(1 βˆ’ r[chosen])
z = (S βˆ’ n) / √n

z > 4 means one-sided p < 3Γ—10⁻⁡. The test works because generator and detector share a deterministic derivation from (key, context) to the exact same pseudo-randomness β€” green lists, random vectors, LSH hyperplanes.

A "same text, with and without watermark" pair cannot exist. In autoregressive generation, changing one word changes everything after it, since the continuation is written on top of that word (in a temperature-0 experiment, both versions matched perfectly through the opening and then split at a single word into an essay and an explainer). What a distortion-free watermark guarantees is not the identity of any individual text but that the distribution of generated texts is unchanged.
03 β€” Fidelity

Fidelity to the paper, and robustness in practice

This implementation carries the full skeleton of the SemStamp paper β€” a key-derived LSH partition of sentence-embedding space, valid regions chained from the previous sentence's signature, sentence-level rejection sampling with a boundary margin, and a z-test on the hit rate. One component was deliberately skipped: the paraphrase-robust encoder.

The paper contrastively fine-tunes the sentence embedder on paraphrase pairs so that rewording never flips an LSH signature. I used off-the-shelf embedders instead (ruri-v3-30m for Japanese, all-MiniLM-L6-v2 for English). The gap is directly measurable as the signature-bit flip rate under paraphrase:

embedderlanguageparaphrase fidelity (cos)LSH bit-flip rate
ruri-v3-30mJapanese0.9524/36 (11%)
all-MiniLM-L6-v2English0.79710/48 (21%)
all-mpnet-base-v2English0.82316/48 (33%)

That flip rate decides everything. All of today's measurements (1–2 runs per condition, 10–13-sentence essays):

conditionwatermarked zafter paraphrase z
Japanese Β· chained regions (paper-faithful)4.024.08 survived
Japanese Β· fixed regions (simplified)4.751.35 / 4.43 (two runs; it varies)
English Β· chained regions2.45-0.82 destroyed
English Β· fixed regions (two runs)4.67 / 4.961.12 / 4.01 (it varies)
(reference) KGW, ja / en7.17 / 14.240.71 / -0.92 destroyed every time

The picture is consistent: KGW sinks to z β‰ˆ 0 after every paraphrase. The semantic watermark degrades gracefully instead of collapsing, and in Japanese it sometimes stays above the detection threshold β€” because ruri's flip rate is only 11%. But as long as flips are non-zero, survival is probabilistic, and 10-sentence texts leave little statistical headroom. The reason the paper bothers to train a robust encoder is precisely to drive that flip rate toward zero β€” turning survival from "sometimes" into "guaranteed".

Two implementation lessons surfaced along the way. (1) Off-the-shelf embeddings are strongly anisotropic, so hyperplanes through the origin partition them unevenly β€” fixed by centering on the mean embedding of a fixed anchor-sentence set. (2) Rejection sampling accepts whatever lands in a valid region regardless of quality, which let garbled or code-switched candidates through β€” fixed by a candidate quality filter applied identically to the watermarked and unwatermarked sides.

04 β€” Limitations

Limitations

05 β€” Future Work

Future work

06 β€” References

References

  1. hellorusk, "Claude started watermarking its text, so I studied how LLM watermarking works" (Japanese), Zenn, 2026. (The article that motivated this project)
  2. Kirchenbauer et al., "A Watermark for Large Language Models" (ICML 2023). β€” KGW
  3. Kuditipudi et al., "Robust Distortion-free Watermarks for Language Models" (2023). β€” Gumbel-Max formalization
  4. Dathathri et al., "Scalable watermarking for identifying large language model outputs" (Nature, 2024). β€” SynthID-Text
  5. Krishna et al., "Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense" (NeurIPS 2023). β€” the paraphrase attack
  6. Hou et al., "SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation" (NAACL 2024). β€” the basis of this implementation, including the contrastively trained robust encoder
  7. Qu et al., "Provably Robust Multi-bit Watermarking for AI-generated Text" (USENIX Security 2025).
  8. Yu et al., "SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling" (NeurIPS 2025).
  9. Ai & He, "PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks" (ICML 2026).
  10. Anthropic, "How Claude marks AI-generated content" (2026).

Models used: sarashina2.2-3b-instruct-v0.1 (SB Intuitions), ruri-v3-30m (Nagoya University), Qwen2.5-instruct 0.5B/1.5B (Alibaba), and all-MiniLM-L6-v2 / all-mpnet-base-v2 (Sentence Transformers) β€” all running locally on Apple Silicon via MPS. Every number on this page was measured on 2026-08-13.