WavePrune: One period is often enough for RoPE
AuthorsGuancheng Du, Luotian Huang, Shaowen Wang, Si Li, Kaifeng Lyu
AffiliationsTsinghua University
WavePrune trims redundant RoPE rotations to improve long-context performance while making attention faster.
Key results
Training-free WavePrune raises the score from 35.7.
Speedup over FlashAttention-2 at 32K context.
Speedup over FlashAttention-2 at 32K context.
With WavePrune in pretraining and evaluation at 4096 tokens; evaluation-only pruning gives 3.3314.
What the paper found
WavePrune targets a subtle problem in Rotary Position Embedding: repeated rotations can make distant token positions look alike, scattering attention onto misleading “ghost” sub-diagonals. The method gives each RoPE channel a local attention window equal to its first rotation period, removing those aliases without fine-tuning. On HELMET, it improves scores for four of five tested models, including Qwen3-8B, which rises from 35.7 to 40.0; it also helps Gemma3-12B and Ministral3-3B, while Llama3.1-8B loses retrieval performance. For efficient execution, FlashWavePrune groups head dimensions in blocks of 16 so its sparse attention computation can use Tensor Cores. Against FlashAttention-2 at 32K context, it delivers 1.15× faster prefill and 1.24× faster decoding on an NVIDIA A100. Training from scratch also improves length extrapolation: on OpenWebText, GPT2-medium trained with WavePrune and evaluated at 4096 tokens reaches validation loss 3.0953, compared with 3.3314 when pruning is added only at evaluation. The results suggest that, for many models, attention beyond RoPE’s first period adds distraction rather than useful context.
Original abstract
Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative positions separated by a full rotation period become hard to tell apart. To address this, we propose WavePrune, which restricts each channel to its first rotation period. We show that it removes the distractions in attention maps created by position aliasing and improves overall long-context performance. Specifically, WavePrune raises the HELMET score on four of five models we test without any extra tuning (e.g., 35.7 -> 40.0 on Qwen3-8B). When pretraining models from scratch, WavePrune also achieves lower validation loss at extrapolated lengths than pretraining without it. Because WavePrune restricts each channel to a sliding window, it induces a fine-grained sparsity that our hardware-aligned CUDA kernels exploit for 1.15x prefill and 1.24x decoding speedups over FlashAttention-2 at 32K context. Together, these results show that RoPE's periodic structure, widely regarded as essential, is largely redundant beyond the first rotation period.
Read the original paperMore in Transformers
Browse all 45 papers →Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers
Neel Varma, Andrew Rufail, Dipika Khullar, Vasu Sharma
This study shows that ViT register tokens carry crucial high-level semantics, while high-norm patch tokens mainly encode lower-level visual structure and contribute little to overall representations.
Length Generalization Needs Proper Regularization
Pavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso
The paper shows that carefully placed dropout can help language models generalize far beyond their training context length.
Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.