NTH

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

AuthorsZhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, Runcong Zhao

June 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that averaging outputs from a few LLMs can wash away text watermarks, exposing a major weakness in current AI-generated text detection methods.

Key results

3
Models evaluated

Experiments were run on three LLMs: Qwen3-8B, Llama-3.1-8B, and Ministral3-8B.

6
Watermarking schemes

The method was tested across six watermarking schemes: AAR, DIPMark, ITS-Edit, KGW, Exp-Edit, and Water-Bag.

5-300
Generation-time z-score range

Watermarked models produced strongly detectable generation-time z-scores in this range, which WASH reduced to below 2.

4
Detection threshold

The paper treats z = 4 as the detection threshold; WASH brings generation-time z-scores below this level.

27.5%
Average quality improvement

WASH improved generation quality by 27.5% on average over the evaluated tasks.

What the paper found

This ICML 2026 paper, “Linear Ensembles Wash Away Watermarks,” by Zhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, and Runcong Zhao shows that distributional watermarking for large language models is fundamentally fragile in the modern multi-provider ecosystem. The authors prove that if independently watermarked models from providers such as Meta’s Llama-3.1-8B, Qwen3-8B, and Ministral3-8B are linearly averaged, the ensemble recovers the unwatermarked consensus distribution with ℓ∞ error shrinking as O(1/√N), up to a second-order variance term. Building on this, they introduce WASH, which combines probability averaging with fluency-aware routing and context re-synchronization to handle mismatched vocabularies and tokenizers across heterogeneous models. Across six watermarking schemes, including AAR, DIPMark, ITS-Edit, KGW, Exp-Edit, and Water-Bag, WASH with just 3 models suppresses generation-time z-scores from 5–300 to below 2, below the detection threshold of 4, and reduces final-text detector TPR@5%FPR to below 50%. It also improves generation quality by 27.5% on average and runs about 6× faster than the strongest removal baselines such as De-mark and ToBlend, which often incur 12×–39× latency overhead. The key novelty is not brute-force rewriting but exploiting the independence of watermark perturbations across competing providers, suggesting that robust provenance detection will require coordinated watermark keys or cross-provider standardization rather than model-level defenses alone.

Original abstract

Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's reality), watermarks trivially fail. Watermarks perturb output distributions away from the original, and in competitive markets, these perturbations are typically independent across providers. We theoretically prove that averaging output probability distributions recovers the unwatermarked distribution with up to a second-order error term. Empirically, simply averaging 3-5 models cancels out these perturbations. We introduce WASH (Watermark Attenuation via Statistical Hybridisation), which solves practical challenges in ensemble generation: vocabulary misalignment and tokenisation differences across heterogeneous models. Experiments across six watermarking schemes and three LLMs show that averaging across 3 models suppresses detection z-scores from 5-300 to below 2 (below the detection threshold of 4) and reduces TPR at 5% FPR to below 50%, while improving quality by 27.5% and running 6 times faster than the best baseline on the long sequence generation. Our results suggest that robust AI-text detection via watermarking requires either accepting this fundamental vulnerability or unprecedented coordination among model providers.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis