Length Generalization Needs Proper Regularization
AuthorsPavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso
AffiliationsInstituto Superior Técnico, Universidade de Lisboa · Instituto de Telecomunicações · ELLIS Unit Lisbon · Gandara AI
The paper shows that carefully placed dropout can help language models generalize far beyond their training context length.
Key results
Pre-Linear dropout retained 100.0 recall-based accuracy at 64× the 4K training context.
Score achieved at 64× the training context with Pre-Linear dropout.
At 44K steps, VPAD outperformed the next-best method by this margin.
Peak RULER score was higher than the weight-decay baseline.
Dropout extended the RULER and HELMET extrapolation range by at least 4×.
What the paper found
Length generalization can deteriorate during training, not just because of positional encodings or attention: models can overfit to the context length they saw in training, and weight decay can accelerate that failure. The paper argues this challenges the strategy of extending context through continued training on longer sequences, used in long-context efforts including NVIDIA’s. It identifies a problem with conventional Pre-Norm dropout: placing dropout before the next normalization layer shifts the mean of activations entering linear projections between training and inference. Moving dropout after normalization, immediately before the linear projection, preserves the mean; the proposed Variance-Preserving Affine Dropout, or VPAD, also matches per-coordinate variance. On SmolLM3, Pre-Linear dropout achieved 100.0 recall-based accuracy on Needle-in-a-Haystack at 64× its 4K training context, while VPAD also extended extrapolation across RULER and HELMET. In scratch-trained transformers, VPAD scored 60% above the next-best method on RULER at 44K steps, and removing weight decay produced a 55% higher peak RULER score than using it. The findings also transfer to Mamba2: dropout extended its RULER and HELMET extrapolation range by at least 4×. Together, the results show that carefully placed dropout can counter length-based overfitting across both transformer and state-space architectures.
Original abstract
Length generalization is the ability of sequential models to perform well on context lengths unseen during training. In this work, we show that the challenge of achieving length generalization is related not only to architectural choices such as positional encoding and the attention mechanism but also to the training procedure itself. We study how regularization affects length generalization and find that weight decay can hinder extrapolation. In contrast, dropout improves extrapolation when its placement within the architecture is reconsidered. In that regard, we show that the standard placement of dropout before layer normalization introduces a systematic distributional mismatch, and that applying dropout just before the linear projection resolves this issue. For example, a modified SmolLM3 with sliding window attention, continually pre-trained with dropout, can extrapolate perfectly to 64$\times$ on Needle-in-a-Haystack and far beyond the pre-training context size on RULER and HELMET. Mamba2 also benefits from dropout, suggesting an architecture-agnostic nature of the problem. We further propose Variance-Preserving Affine Dropout (VPAD), a new dropout strategy that substantially reduces the resulting pre-activation variance mismatch, leading to further extrapolation improvements in transformers.
Read the original paperMore in Transformers
Browse all 45 papers →Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers
Neel Varma, Andrew Rufail, Dipika Khullar, Vasu Sharma
This study shows that ViT register tokens carry crucial high-level semantics, while high-norm patch tokens mainly encode lower-level visual structure and contribute little to overall representations.
WavePrune: One period is often enough for RoPE
Guancheng Du, Luotian Huang, Shaowen Wang, Si Li, Kaifeng Lyu
WavePrune trims redundant RoPE rotations to improve long-context performance while making attention faster.
Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.