Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
AuthorsZhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, Runcong Zhao
Resources
This paper shows that averaging outputs from a few LLMs can wash away text watermarks, exposing a major weakness in current AI-generated text detection methods.
Key results
Experiments were run on three LLMs: Qwen3-8B, Llama-3.1-8B, and Ministral3-8B.
The method was tested across six watermarking schemes: AAR, DIPMark, ITS-Edit, KGW, Exp-Edit, and Water-Bag.
Watermarked models produced strongly detectable generation-time z-scores in this range, which WASH reduced to below 2.
The paper treats z = 4 as the detection threshold; WASH brings generation-time z-scores below this level.
WASH improved generation quality by 27.5% on average over the evaluated tasks.
What the paper found
This ICML 2026 paper, “Linear Ensembles Wash Away Watermarks,” by Zhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, and Runcong Zhao shows that distributional watermarking for large language models is fundamentally fragile in the modern multi-provider ecosystem. The authors prove that if independently watermarked models from providers such as Meta’s Llama-3.1-8B, Qwen3-8B, and Ministral3-8B are linearly averaged, the ensemble recovers the unwatermarked consensus distribution with ℓ∞ error shrinking as O(1/√N), up to a second-order variance term. Building on this, they introduce WASH, which combines probability averaging with fluency-aware routing and context re-synchronization to handle mismatched vocabularies and tokenizers across heterogeneous models. Across six watermarking schemes, including AAR, DIPMark, ITS-Edit, KGW, Exp-Edit, and Water-Bag, WASH with just 3 models suppresses generation-time z-scores from 5–300 to below 2, below the detection threshold of 4, and reduces final-text detector TPR@5%FPR to below 50%. It also improves generation quality by 27.5% on average and runs about 6× faster than the strongest removal baselines such as De-mark and ToBlend, which often incur 12×–39× latency overhead. The key novelty is not brute-force rewriting but exploiting the independence of watermark perturbations across competing providers, suggesting that robust provenance detection will require coordinated watermark keys or cross-provider standardization rather than model-level defenses alone.
Original abstract
Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's reality), watermarks trivially fail. Watermarks perturb output distributions away from the original, and in competitive markets, these perturbations are typically independent across providers. We theoretically prove that averaging output probability distributions recovers the unwatermarked distribution with up to a second-order error term. Empirically, simply averaging 3-5 models cancels out these perturbations. We introduce WASH (Watermark Attenuation via Statistical Hybridisation), which solves practical challenges in ensemble generation: vocabulary misalignment and tokenisation differences across heterogeneous models. Experiments across six watermarking schemes and three LLMs show that averaging across 3 models suppresses detection z-scores from 5-300 to below 2 (below the detection threshold of 4) and reduces TPR at 5% FPR to below 50%, while improving quality by 27.5% and running 6 times faster than the best baseline on the long sequence generation. Our results suggest that robust AI-text detection via watermarking requires either accepting this fundamental vulnerability or unprecedented coordination among model providers.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.