AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
AuthorsSaifur Rahman Tamim, Amir Labib Khan
Resources
The study finds that current LLM watermarks are too fragile and error-prone to serve as reliable courtroom evidence, especially after meaning-preserving paraphrase.
Key results
Combined valid runs across KGW, Unigram, and SynthID.
Initially detected KGW watermarks removed after meaning-preserving paraphrase.
Initially detected SynthID watermarks removed after paraphrase.
Paraphrased human-written controls falsely classified as AI-generated.
Pristine SynthID outputs placed in the detector’s uncertainty deadband.
What the paper found
This study asks whether LLM watermark detections can survive courtroom scrutiny under the Daubert standard and NIST SP 800-86 forensic workflow. Using the MarkLLM toolkit, researchers tested KGW, Unigram, and the open-source MarkLLM implementation of Google’s SynthID-Text, generated with Qwen2.5-1.5B or Gemma-2-9b-it and attacked with meaning-preserving paraphrases from Qwen2.5-1.5B. Across 846 valid paraphrase runs, semantic similarity remained high, with retained outputs above a cosine threshold of 0.75, yet every initially detected KGW and Unigram sample lost its watermark—100% conditional removal—while SynthID lost detection in 98.3% of cases. Detection was already weak before attack: false-negative rates were 70% for KGW, 83% for Unigram, and 80% for SynthID. SynthID also falsely flagged 5.4% of paraphrased human-written controls and placed 80% of its pristine watermarked outputs in an uncertainty deadband. The authors propose a Forensic Readiness Score with 12 criteria, three mandatory gates, and a 60-point scale, but show that point systems can conceal practical failure: Unigram reaches the conditional threshold despite the worst false-negative rate and total conditional removal. All three methods fail at least two of five Daubert factors, especially stable error rates and controlling standards. The paper’s conclusion is not that watermark detectors are random; their outputs are repeatable, but meaning-preserving edits reproducibly destroy the evidentiary signal, making current configurations unsuitable as courtroom-grade proof.
Original abstract
Governments are increasingly mandating that LLM-generated content carry watermarks. The EU AI Act calls for markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that is "permanent or extraordinarily difficult to remove." Both mandates rest on an untested assumption: that watermark detection yields evidence reliable enough for courts. This paper tests that assumption directly. We evaluate three representative LLM watermarking methods -- KGW, Unigram, and the MarkLLM implementation of SynthID-Text -- against the Daubert admissibility criteria and the NIST SP 800-86 digital forensic process. To structure this evaluation, we propose a Forensic Readiness Score (FRS) framework with 12 criteria, three mandatory gates, and a 60-point scoring system. We focus on meaning-preserving paraphrase as the attack vector, since it is both legally realistic and difficult to dismiss as evidence tampering. The results raise serious evidentiary concerns. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing -- 100% conditional removal. SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfy more than two of five Daubert factors. We also find that the FRS point-based scoring system, despite working as designed, cannot fully capture forensic uselessness -- a limitation worth noting for future framework design. These configurations, as tested, do not meet the evidentiary bar that courts require.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.