Paper reading on Kirchenbauer et al. 2023
An AI text watermark is not a hidden character, a secret tag, or special file metadata. Instead, text watermarks are mathematical patterns in word choice
A brief explanation how LLM works
A brief idea of AI watermark
Before generating word: The AI looks at the word it just typed and uses a secret rule to secretly split the entire dictionary into 50% Green words and 50% Red words.
During generating words: The AI gives Green words a subtle extra boost (\(\delta\)), making them far more likely to be chosen while keeping the sentence completely natural.
Detecting a text: A human naturally lands on Green about 50% of the time (like flipping a fair coin), but the watermarked AI lands on Green 75%–85% of the time.
The final judgement: If a 200-word paragraph hits Green far too often to be pure luck, a quick probability test proves it came from the AI.
Assuming a 50% (\(\gamma\)) Green/Red split
The Expectation of Z-test statistics given a sample of T tokens is: \[ \mathbb{E}[z] = \frac{(\bar{p}_G - \gamma)\sqrt{T}}{\sqrt{\gamma(1 - \gamma)}} \]
\(\bar{p}_G\) is the observed proportion of Green words in the sample, this rate is roughly controlled by how you set the watermarking strength (\(\delta\)).
In common settting we use the critial value of 3 (\(\tau\)) to determine if a text is watermarked or not.
We can plot how does \(\bar{p}_G\) vary with different sample sizes.
| Watermark Strength | Expected Green Rate (\(\bar{p}_G\)) | Minimum Detectable Sample Size (T) | Minimum Detectable Sample Size (Words) |
|---|---|---|---|
| Moderate Soft Watermark (\(\delta \approx 1.5\)) | \(\bar{p}_G \approx 0.70\) | \(\approx 100\) tokens (~75 words) | \(\approx 56\) tokens |
| Standard Soft Watermark (\(\delta \approx 2.0\)) | \(\bar{p}_G \approx 0.80\) | \(\approx 44\) tokens (~30–35 words) | \(\approx 25\) tokens |
| Aggressive Soft Watermark (\(\delta \ge 4.0\)) | \(\bar{p}_G \approx 0.90\) | \(\approx 25\) tokens (~18–20 words) | \(\approx 14\) tokens |
| Theoretical Limit: Hard Red List (\(\delta \to \infty\)) | \(\bar{p}_G = 1.0\) (100% Green) | \(16\) tokens | \(9\) tokens |
| Sequence Length (T) | False Negative Rate (\(\bar{p}_G \approx 0.70\)) | False Negative Rate (\(\bar{p}_G \approx 0.80\)) | False Negative Rate (\(\bar{p}_G \approx 0.90\)) |
|---|---|---|---|
| \(T = 25\) tokens (~18 words) | \(\approx 99.8\%\) (undetermined) | \(\approx 84.1%\) | \(\approx 15.9\%\) |
| \(T = 50\) tokens | \(\approx 92.3\%\) | \(\approx 10.6\%\) | \(\approx 0.02\%\) |
| \(T = 100\) tokens | \(\approx 50.0\%\) | \(\approx 0.03\%\) | \(< 10^{-6}\) |
| \(T = 200\) tokens | \(\approx 1.4\%\) | \(< 10^{-6}\) | \(\approx 0\) |
| \(T \ge 300\) tokens | \(\approx 0.003\%\) | \(\approx 0\) | \(\approx 0\) |
LLM1 -> “Watermarked” -> LLM2 -> “Unwatermarked”
Usually LLM2 will be some open-source model that is not watermarked, and it will paraphrase the text in a way that removes the watermark.
Laziest approach to avoid detection, but it can also introduce errors or change the meaning of the text.
| Scenario | Watermark Risk | Real-World Impact |
|---|---|---|
| Brainstorming & Outlining | Zero | Ideas, bullet points, and high-level concepts carry no sentence-level token distribution. |
| Proofreading / Polishing your own draft | Negligible | If the ideas and majority of words are yours, the AI merely fixes grammar or word order. There is very little AI token sequence for a detector to latch onto. |
| Direct Generation (Copy-Pasting full drafts) | High | The text contains raw, continuous model token sequences that reliably trigger detection algorithms. |
Text watermarking is not the only modality that can be watermarked. Other modalities include:
Still text watermarking is more challenging due to the discrete nature of language and the need for semantic coherence.