NTH

WavePrune: One period is often enough for RoPE

AuthorsGuancheng Du, Luotian Huang, Shaowen Wang, Si Li, Kaifeng Lyu

AffiliationsTsinghua University

October 10, 2026 2 min read
Watch on YouTube
The one-line take

WavePrune trims redundant RoPE rotations to improve long-context performance while making attention faster.

Key results

40.0
Qwen3-8B HELMET score

Training-free WavePrune raises the score from 35.7.

1.15
Prefill speedup

Speedup over FlashAttention-2 at 32K context.

1.24
Decoding speedup

Speedup over FlashAttention-2 at 32K context.

3.0953
GPT2-medium extrapolated validation loss

With WavePrune in pretraining and evaluation at 4096 tokens; evaluation-only pruning gives 3.3314.

What the paper found

WavePrune targets a subtle problem in Rotary Position Embedding: repeated rotations can make distant token positions look alike, scattering attention onto misleading “ghost” sub-diagonals. The method gives each RoPE channel a local attention window equal to its first rotation period, removing those aliases without fine-tuning. On HELMET, it improves scores for four of five tested models, including Qwen3-8B, which rises from 35.7 to 40.0; it also helps Gemma3-12B and Ministral3-3B, while Llama3.1-8B loses retrieval performance. For efficient execution, FlashWavePrune groups head dimensions in blocks of 16 so its sparse attention computation can use Tensor Cores. Against FlashAttention-2 at 32K context, it delivers 1.15× faster prefill and 1.24× faster decoding on an NVIDIA A100. Training from scratch also improves length extrapolation: on OpenWebText, GPT2-medium trained with WavePrune and evaluated at 4096 tokens reaches validation loss 3.0953, compared with 3.3314 when pruning is added only at evaluation. The results suggest that, for many models, attention beyond RoPE’s first period adds distraction rather than useful context.

Original abstract

Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative positions separated by a full rotation period become hard to tell apart. To address this, we propose WavePrune, which restricts each channel to its first rotation period. We show that it removes the distractions in attention maps created by position aliasing and improves overall long-context performance. Specifically, WavePrune raises the HELMET score on four of five models we test without any extra tuning (e.g., 35.7 -> 40.0 on Qwen3-8B). When pretraining models from scratch, WavePrune also achieves lower validation loss at extrapolated lengths than pretraining without it. Because WavePrune restricts each channel to a sliding window, it induces a fine-grained sparsity that our hardware-aligned CUDA kernels exploit for 1.15x prefill and 1.24x decoding speedups over FlashAttention-2 at 32K context. Together, these results show that RoPE's periodic structure, widely regarded as essential, is largely redundant beyond the first rotation period.

Read the original paper

More in Transformers

Browse all 45 papers →
02Transformer

Length Generalization Needs Proper Regularization

Pavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso

The paper shows that carefully placed dropout can help language models generalize far beyond their training context length.

Read analysis