NTH

Length Generalization Needs Proper Regularization

AuthorsPavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso

AffiliationsInstituto Superior Técnico, Universidade de Lisboa · Instituto de Telecomunicações · ELLIS Unit Lisbon · Gandara AI

October 10, 2026 2 min read
Watch on YouTube
The one-line take

The paper shows that carefully placed dropout can help language models generalize far beyond their training context length.

Key results

64
SmolLM3 NIAH extrapolation

Pre-Linear dropout retained 100.0 recall-based accuracy at 64× the 4K training context.

100.0
SmolLM3 NIAH recall-based accuracy

Score achieved at 64× the training context with Pre-Linear dropout.

60%
VPAD RULER improvement

At 44K steps, VPAD outperformed the next-best method by this margin.

55%
No-weight-decay RULER improvement

Peak RULER score was higher than the weight-decay baseline.

4
Mamba2 extrapolation range

Dropout extended the RULER and HELMET extrapolation range by at least 4×.

What the paper found

Length generalization can deteriorate during training, not just because of positional encodings or attention: models can overfit to the context length they saw in training, and weight decay can accelerate that failure. The paper argues this challenges the strategy of extending context through continued training on longer sequences, used in long-context efforts including NVIDIA’s. It identifies a problem with conventional Pre-Norm dropout: placing dropout before the next normalization layer shifts the mean of activations entering linear projections between training and inference. Moving dropout after normalization, immediately before the linear projection, preserves the mean; the proposed Variance-Preserving Affine Dropout, or VPAD, also matches per-coordinate variance. On SmolLM3, Pre-Linear dropout achieved 100.0 recall-based accuracy on Needle-in-a-Haystack at 64× its 4K training context, while VPAD also extended extrapolation across RULER and HELMET. In scratch-trained transformers, VPAD scored 60% above the next-best method on RULER at 44K steps, and removing weight decay produced a 55% higher peak RULER score than using it. The findings also transfer to Mamba2: dropout extended its RULER and HELMET extrapolation range by at least 4×. Together, the results show that carefully placed dropout can counter length-based overfitting across both transformer and state-space architectures.

Original abstract

Length generalization is the ability of sequential models to perform well on context lengths unseen during training. In this work, we show that the challenge of achieving length generalization is related not only to architectural choices such as positional encoding and the attention mechanism but also to the training procedure itself. We study how regularization affects length generalization and find that weight decay can hinder extrapolation. In contrast, dropout improves extrapolation when its placement within the architecture is reconsidered. In that regard, we show that the standard placement of dropout before layer normalization introduces a systematic distributional mismatch, and that applying dropout just before the linear projection resolves this issue. For example, a modified SmolLM3 with sliding window attention, continually pre-trained with dropout, can extrapolate perfectly to 64$\times$ on Needle-in-a-Haystack and far beyond the pre-training context size on RULER and HELMET. Mamba2 also benefits from dropout, suggesting an architecture-agnostic nature of the problem. We further propose Variance-Preserving Affine Dropout (VPAD), a new dropout strategy that substantially reduces the resulting pre-activation variance mismatch, leading to further extrapolation improvements in transformers.

Read the original paper

More in Transformers

Browse all 45 papers →
02Transformer

WavePrune: One period is often enough for RoPE

Guancheng Du, Luotian Huang, Shaowen Wang, Si Li, Kaifeng Lyu

WavePrune trims redundant RoPE rotations to improve long-context performance while making attention faster.

Read analysis