Distance generalization in transformers: why bother with positional encoding?
AuthorsDaniel Henrik Nevermann, Claudius Gros
Resources
This paper tests whether transformers really need positional encodings to recognize token distances, revealing when different schemes help or hurt generalization to unseen delays.
Key results
Fixed context window used for most training and evaluation.
Causal decoder-only transformer depth.
Transformer model dimension.
Lower endpoint of the primary training delay interval, which spans 15 to 25.
Approximate dataset size used for training.
What the paper found
This paper isolates distance generalization: whether a transformer can copy tokens when the source-to-recall delay changes, while the overall context window stays fixed, separating unseen token dependencies from unseen positions in length generalization. Experiments use full and selective delay-copy tasks with a causal decoder-only transformer containing 8 layers, 8 attention heads, and hidden size 512, trained on 10M synthetic tokens with a fixed context length of 256. Models trained on delays from 15 to 25 show a counterintuitive ranking: NoPE, with no explicit positional encoding, usually generalizes better to unseen distances than ALiBi and RoPE, even though RoPE is widely used in models such as Llama 3 and DeepSeek-V3. The result suggests causal attention can infer relative position dynamically, although NoPE becomes unreliable in smaller models. Increasing training diversity across delay distances improves absolute out-of-distribution accuracy, but with diminishing relative returns. Transfer learning between full and selective copying is conditional: NoPE and ALiBi often benefit at moderate extrapolation distances but can suffer at larger ones, while RoPE generally exhibits weaker transfer and, in some settings, negative interference. The study argues that distance generalization should be evaluated separately from context-length extrapolation and that positional encoding choices require reconsideration.
Original abstract
Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distances between source and recall, where tokens are copied either fully or selectively, and test models on delays unseen during training. We address three questions: (A) Do positional encoding schemes such as RoPE and ALiBi improve distance resolution relative to no positional encoding (NoPE)? (B) How does data diversity, the number of inter-token distances seen in training, affect performance? (C) When is distance transfer learning positive or negative? We present a thorough investigation, finding that it is paramount to improve our understanding of the underlying mechanisms.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.