In three lines
The watermark lives in the interchangeable choices an LLM makes when picking the next token β places where either word would be natural. Each single choice reveals nothing, but summed over hundreds of tokens the key-driven bias produces a statistical signature (a z-score) that human text cannot exhibit. Token-level watermarks reliably die after one paraphrase; the semantic-space watermark degrades but can survive β and its survival hinges on how paraphrase-stable the sentence embedder is.
Results: watermarked and unwatermarked, side by side
Example 1 β Semantic watermark (SemStamp-style): degrades, but does not collapse
To make the pair directly comparable, the prompt pins the structure: both essays must open with the exact sentence "Why do people enjoy walking in the rain?" and cover, in order, the sound of rain, the smell and the air, and the calming effect on the mind. Generated in English by Qwen2.5-0.5B-instruct with all-MiniLM-L6-v2 sentence embeddings (fixed-region variant; see the fidelity section for the chained variant). Within that shared skeleton, the wording differs because sampling differs on every run β not because of the watermark. The watermark leaves no visible trace; it exists only in the statistic a key-holding detector can compute. A 0.5B model writes rough English; the texts are shown unedited.
"Why do people enjoy walking in the rain?" is a question that speaks to our senses, emotions, and experiences. For deciding when to go for a walk during rainy weather is a personal decision based on various factors such as the climate, personal preferences, or even the weather forecast. Firstly, there's the auditory aspect of rainfall, which is often more prevalent than visual. The rhythmic thumping of waterfalls creates a soothing background noise that can help in calming oneβs mind, making it easier to relax during a rainy day. Raindrops falling on wet surfaces create a sense of moisture and humidity, which can contribute to feelings of calmness and tranquility. Secondly, the tactile experience of walking in the rain adds an extra layer of sensory stimulation. The ground underfoot feels like carpeted grass beneath your feet, adding a touch of serenity to the already pleasant texture of wet earth. This tactile engagement also helps to increase body awareness and connection with nature around you. Lastly, the mental benefits associated with enjoying outdoor activities are equally significant. The feeling of being immersed in fresh air, surrounded by lush greenery, and feeling lighter overall during rainy days leaves many people feeling refreshed and rejuvenated.
Why Do People Enjoy Walking in the Rain? People often seek out activities that offer fresh, natural environments to unwind amidst the hustle and bustle of urban life. Among these serene pursuits is one that has garnered widespread appealβwalking through the rain. The allure of nature's embrace of the elements, combined with its ability to bring a refreshing freshness, draws countless individuals into this outdoor activity. One of the primary reasons why walking in the rain evokes such fond memories is its captivating combination of sound and sensory experience. The rhythmic tapping of water as it falls against your feet can be a soothing lullaby, echoing across the landscape and providing a gentle yet constant melody that enhances the tranquility of the moment. This auditory cue creates an atmosphere of peace and calmness, allowing for introspection and a sense of well-being. Moreover, the smell of wet earth is another delightful aspect of walking in the rain. As moisture seeps through the soles of your shoes, the scent of damp grass and leaves fills the air, making you feel refreshed and rejuvenated. This pleasant aroma complements the sounds around you, creating a harmonious ambiance that encourages a relaxed state of mind. The air itself becomes particularly comforting during rainy days. The misty drizzle blankets the landscape, painting everything in a mellow hue of verdant greens and soft pinks. These natural elements not only enhance the aesthetic appeal but also provide a soothing contrast to the harsh exterior of modern cities.
I then had an unwatermarked model (the same LLM with watermarking switched off) rewrite the watermarked essay sentence by sentence, keeping each sentence's meaning β simulating an attacker laundering the text through a clean LLM, which adds no watermark of its own. The wording changed substantially, and β
Why do folks tend to enjoy wandering down below while it rains? People often enjoy pursuing activities that place a strong emphasis on finding environments free from modern noise and chaos, such as forests, vineyards, or even waking up early and venturing out into the evening fields. Among these idyllic activities is one that has gained widespread loveβwalking amidst the drizzle. The allure of nature's embrace of the elements and its capacity to refresh one quickly becomes irresistible, attracting and inspiring countless individuals to pursue this outdoor activity. A primary reason why walking in the rain evokes cherished memories is its striking blend of auditory and sensory delights. The rhythmic tapping of water as it falls against your feet can be a soothing lullaby. It echoes across the landscape, and it provides a gentle yet constant melody that enhances the moment's tranquility. This auditory stimulation produces a tranquil ambiance, promoting mindfulness and tranquility. In addition, the fragrant odors of damp soil are another delightful component of exploring in rain. When moisture seeps through the soles of your shoes, a pleasant aroma of damp grass and leaves fills the air, making you feel revitalized and refreshed. This pleasant aroma complements the sounds around you, elevating the mood and inducing a serene state of mind. On rainy days, the air feels particularly comforting. The drizzle-covered landscape descends in a palette of verdant greens and soft pinks. These natural elements not only grace the aesthetic decor but also serve as a natural counterpoint to the harsh exterior urban landscapes.
Example 2 β Token-space watermark (KGW): detected strongly, erased reliably
The same English prompt watermarked with KGW (a logit bonus on a key-derived "green list" of tokens, Ξ΄ = 2). The detector calls it with overwhelming confidence β and the same paraphrase attack erases it completely, in every trial. Note the prose itself is noticeably degraded ("cuisinarily encode", "the glands of your glands"): on a small model, the logit bonus visibly distorts generation, which is exactly the quality cost the distortion-free schemes were invented to avoid.
In damp, warm weather, there is no glory of walking like the soft embrace of the rain. It is the infinitude of the embrace that is unparalleled in nature. Raindrops falling from astronomical freedom are also serene. Rain is a mystery, sublime, mysterious.
Once in the headlands and orienteering coordinates of a rain forest or a quad of rain forests entwine cuisinarily encode the senses of the individual as if existing out of and from the mediocre fusion of public consciousness. Rain is intimately nurtured by our natural existence. Now, would spell βRainβs the only joy.β Itβs the interpersonal nurture that we bloom with saturated in tenacity, that feels unparalleled in the social jungle where one walks too hard in a coffee-trough under a rain in the beauty of the substance.
Rain, the savor is without a doubt non-slip. The very brightness of the sun as it sheds a gentle and incomprehending embrace across them is accompanied by the powerful mix of fragrances and chemicals filled into an amber-stained sweat of your skin that spreads from the glands of your glands into your body. (β¦)
Example 3 β Where does the watermark live? Coloring every token
If two texts look indistinguishable, where is the watermark? At temperature 0 (deterministic decoding, no sampling noise) both versions even begin with the same eight tokens β "Walking in the rain can be a delightful" β before the logit bonus flips one close call and the texts fork. Running the KGW detector's own computation and coloring each token by whether the key favors it ("green") makes the watermark visible: the unwatermarked text is 20.0% green, near the chance level, while the watermarked one is 55.0% (Qwen2.5-0.5B-instruct, same prompt, temperature 0).
One of the reasons people enjoy walking in the rain is that it can be a great way to explore new places and discover new things. The rain can create a sense of mystery and wonder, and it can be a great opportunity to take in the sights and sounds of the world around you. Walking in the rain can also be a way to connect with nature and appreciate the beauty of the world around us.
Another reason people enjoy walking in the rain is that it can be a way to relax and unwind. The rain can create a sense of calm and tranquility, and it can be a great way to take a break from the stresses of everyday life.
Walking in the rain can also be a refreshing experience, especially if it is a warm day with a light breeze. The raindrops falling onto the grass and pavement can provide a refreshing contrast to the warm, humid air, and the sound of the rain can create a gentle rhythm that is soothing to the ears.
In addition to its calming effects, walking in the rain can also be a great way to explore the outdoors and get some fresh air. The gentle breeze and the sound of the rain can create a peaceful and invigorating atmosphere, making it a perfect activity for a relaxing and enjoyable walk.
Overall, people who enjoy walking in the rain find
No individual green token is special β "rain" and "walking" can be green or red depending on the key and the preceding four tokens at that position. The difference exists only in the aggregate over hundreds of tokens.
Example 4 β Quantitative experiments (both token-space schemes)
Detection z-scores with Qwen2.5-0.5B-instruct, averaged over 4 prompts Γ 200 tokens. Only watermarked text is detected; unwatermarked output, human text (the opening of Austen's Pride and Prejudice), and detection with a wrong key all pass.
| text under test | KGW detector z | Gumbel detector z |
|---|---|---|
| no watermark (vanilla) | -0.76 | 0.73 |
| KGW-watermarked | 9.78 | -0.07 |
| Gumbel-Max-watermarked | -0.90 | 34.99 |
| wrong key | -0.42 | 0.16 |
| human text (Austen) | -1.51 | 0.18 |
Attack experiments: random token substitution degrades the signal gradually, and a single paraphrase pass with an unwatermarked LLM destroys both schemes.
| attack | KGW z | Gumbel z |
|---|---|---|
| none | 5.55 | 38.79 |
| substitution 10% | 2.45 | 19.61 |
| substitution 30% | -0.82 | 7.27 |
| substitution 50% | 1.31 | 0.40 |
| paraphrase (unwatermarked LLM) | -1.45 | -0.66 |
On quality: KGW worsened mean NLL under the same model from 2.25 to 3.02 nats/token (a distorted distribution), while Gumbel-Max stayed at 2.77 β no degradation, exactly as the distortion-free construction promises.
How it works
An LLM samples its next token from a probability distribution, and a passage offers many moments where "either this or that reads naturally". A watermark nudges these interchangeable choices with a rule derived from a secret key. One choice reveals nothing; hundreds of them, aggregated, reveal a statistically impossible habit. Detection needs no model access β only the key and the text.
KGW
Hash the preceding tokens with the key to split the vocabulary into a green list (25%) and a red list, then add +Ξ΄ to green logits. Detection is a z-test on the green-token rate. Simple, but it distorts the distribution and can hurt quality.
early implementation β preserved in git historyGumbel-Max
Replace the sampling randomness with key-derived randoms r and pick argmax ri1/pi. The output distribution is provably unchanged β zero quality cost. Detection scores Ξ£ βlog(1βr). SynthID-Text belongs to this family.
early implementation β preserved in git historySemStamp-style (semantic)
Generate sentence by sentence, accepting only sentences whose embedding lands in "valid regions" of a key-derived LSH partition, with the valid set chained from the previous sentence's signature. Paraphrasing barely moves a sentence's meaning vector, so the watermark can survive.
what sukashi is nowEvery detector is a hypothesis test
All three schemes reduce detection to one question: how many Ο does this text's statistic deviate from the no-watermark null? KGW and SemStamp use the normal approximation of a binomial test; Gumbel uses a sum of exponentials (Gamma null).
# KGW / SemStamp: green (or valid-region) rate test
z = (hits β Ξ³Β·n) / β(nΒ·Ξ³Β·(1βΞ³))
# Gumbel-Max: per-token score is Exp(1) under the null
S = Ξ£ βlog(1 β r[chosen])
z = (S β n) / βn
z > 4 means one-sided p < 3Γ10β»β΅. The test works because generator and detector share a deterministic derivation from (key, context) to the exact same pseudo-randomness β green lists, random vectors, LSH hyperplanes.
Fidelity to the paper, and robustness in practice
This implementation carries the full skeleton of the SemStamp paper β a key-derived LSH partition of sentence-embedding space, valid regions chained from the previous sentence's signature, sentence-level rejection sampling with a boundary margin, and a z-test on the hit rate. One component was deliberately skipped: the paraphrase-robust encoder.
The paper contrastively fine-tunes the sentence embedder on paraphrase pairs so that rewording never flips an LSH signature. I used off-the-shelf embedders instead (ruri-v3-30m for Japanese, all-MiniLM-L6-v2 for English). The gap is directly measurable as the signature-bit flip rate under paraphrase:
| embedder | language | paraphrase fidelity (cos) | LSH bit-flip rate |
|---|---|---|---|
| ruri-v3-30m | Japanese | 0.952 | 4/36 (11%) |
| all-MiniLM-L6-v2 | English | 0.797 | 10/48 (21%) |
| all-mpnet-base-v2 | English | 0.823 | 16/48 (33%) |
That flip rate decides everything. All of today's measurements (1β2 runs per condition, 10β13-sentence essays):
| condition | watermarked z | after paraphrase z |
|---|---|---|
| Japanese Β· chained regions (paper-faithful) | 4.02 | 4.08 survived |
| Japanese Β· fixed regions (simplified) | 4.75 | 1.35 / 4.43 (two runs; it varies) |
| English Β· chained regions | 2.45 | -0.82 destroyed |
| English Β· fixed regions (two runs) | 4.67 / 4.96 | 1.12 / 4.01 (it varies) |
| (reference) KGW, ja / en | 7.17 / 14.24 | 0.71 / -0.92 destroyed every time |
The picture is consistent: KGW sinks to z β 0 after every paraphrase. The semantic watermark degrades gracefully instead of collapsing, and in Japanese it sometimes stays above the detection threshold β because ruri's flip rate is only 11%. But as long as flips are non-zero, survival is probabilistic, and 10-sentence texts leave little statistical headroom. The reason the paper bothers to train a robust encoder is precisely to drive that flip rate toward zero β turning survival from "sometimes" into "guaranteed".
Two implementation lessons surfaced along the way. (1) Off-the-shelf embeddings are strongly anisotropic, so hyperplanes through the origin partition them unevenly β fixed by centering on the mean embedding of a fixed anchor-sentence set. (2) Rejection sampling accepts whatever lands in a valid region regardless of quality, which let garbled or code-switched candidates through β fixed by a candidate quality filter applied identically to the watermarked and unwatermarked sides.
Limitations
- Short texts cannot be detected. Evidence accumulates with token (or sentence) count. Empirically, a 57-token constrained list scored z = 3.90 β below threshold β while the 113-token version of the same format reached z = 5.38. The sentence-level semantic scheme needs even longer texts.
- Low-entropy passages resist watermarking. Code or factual statements with essentially one correct continuation offer no interchangeable choices to bias. Forcing KGW onto them surfaces as quality damage (measured: the Ξ΄ = 2 logit bonus broke formatting-instruction compliance and caused code-switching).
- Semantic robustness depends on the embedder. With off-the-shelf embedders, paraphrase survival is probabilistic (see the table above). Translation, summarization, or heavy rewriting that reorganizes sentence-level meaning defeats the scheme even with a robust encoder.
- Detection is not proof of AI authorship. Having an LLM proofread a human's text also imprints the watermark. Detection shows only that the model (key) processed the text β Anthropic states the same caveat about Claude's watermark.
- This is a minimal educational implementation. Robust-encoder training, key rotation, multiple keys, and production-scale false-positive evaluation (SynthID-Text was evaluated on ~20M live requests) are out of scope.
Future work
- Training the paraphrase-robust encoder β the one missing component. For Japanese, contrastively fine-tuning ruri on paraphrase pairs (JSNLI, PAWS-X) should push the flip rate to a few percent, turning survival from occasional into reliable.
- Multi-bit payloads. Embedding and decoding user IDs or timestamps rather than a single "AI or not" bit (Qu et al., USENIX Security 2025) β the leading direction for content source tracing.
- Post-hoc watermarking. Watermarking through a black-box API with no logit access (SAEMark, NeurIPS 2025), selecting among candidate outputs by Sparse Autoencoder features.
- Unifying distortion-freeness with paraphrase robustness (PASA, ICML 2026) β the two properties this project verified separately, achieved in one principled framework.
- Checking against Anthropic's detector. The technical details of Claude's watermark are unpublished; once a detection tool ships, I want to compare its behavior with this implementation.
References
- hellorusk, "Claude started watermarking its text, so I studied how LLM watermarking works" (Japanese), Zenn, 2026. (The article that motivated this project)
- Kirchenbauer et al., "A Watermark for Large Language Models" (ICML 2023). β KGW
- Kuditipudi et al., "Robust Distortion-free Watermarks for Language Models" (2023). β Gumbel-Max formalization
- Dathathri et al., "Scalable watermarking for identifying large language model outputs" (Nature, 2024). β SynthID-Text
- Krishna et al., "Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense" (NeurIPS 2023). β the paraphrase attack
- Hou et al., "SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation" (NAACL 2024). β the basis of this implementation, including the contrastively trained robust encoder
- Qu et al., "Provably Robust Multi-bit Watermarking for AI-generated Text" (USENIX Security 2025).
- Yu et al., "SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling" (NeurIPS 2025).
- Ai & He, "PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks" (ICML 2026).
- Anthropic, "How Claude marks AI-generated content" (2026).
Models used: sarashina2.2-3b-instruct-v0.1 (SB Intuitions), ruri-v3-30m (Nagoya University), Qwen2.5-instruct 0.5B/1.5B (Alibaba), and all-MiniLM-L6-v2 / all-mpnet-base-v2 (Sentence Transformers) β all running locally on Apple Silicon via MPS. Every number on this page was measured on 2026-08-13.