NTH

Distance generalization in transformers: why bother with positional encoding?

AuthorsDaniel Henrik Nevermann, Claudius Gros

September 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper tests whether transformers really need positional encodings to recognize token distances, revealing when different schemes help or hurt generalization to unseen delays.

Key results

256
Context length

Fixed context window used for most training and evaluation.

8
Transformer layers

Causal decoder-only transformer depth.

512
Hidden size

Transformer model dimension.

15
Training delay lower bound

Lower endpoint of the primary training delay interval, which spans 15 to 25.

10M
Synthetic training tokens

Approximate dataset size used for training.

What the paper found

This paper isolates distance generalization: whether a transformer can copy tokens when the source-to-recall delay changes, while the overall context window stays fixed, separating unseen token dependencies from unseen positions in length generalization. Experiments use full and selective delay-copy tasks with a causal decoder-only transformer containing 8 layers, 8 attention heads, and hidden size 512, trained on 10M synthetic tokens with a fixed context length of 256. Models trained on delays from 15 to 25 show a counterintuitive ranking: NoPE, with no explicit positional encoding, usually generalizes better to unseen distances than ALiBi and RoPE, even though RoPE is widely used in models such as Llama 3 and DeepSeek-V3. The result suggests causal attention can infer relative position dynamically, although NoPE becomes unreliable in smaller models. Increasing training diversity across delay distances improves absolute out-of-distribution accuracy, but with diminishing relative returns. Transfer learning between full and selective copying is conditional: NoPE and ALiBi often benefit at moderate extrapolation distances but can suffer at larger ones, while RoPE generally exhibits weaker transfer and, in some settings, negative interference. The study argues that distance generalization should be evaluated separately from context-length extrapolation and that positional encoding choices require reconsideration.

Original abstract

Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distances between source and recall, where tokens are copied either fully or selectively, and test models on delays unseen during training. We address three questions: (A) Do positional encoding schemes such as RoPE and ALiBi improve distance resolution relative to no positional encoding (NoPE)? (B) How does data diversity, the number of inter-token distances seen in training, affect performance? (C) When is distance transfer learning positive or negative? We present a thorough investigation, finding that it is paramount to improve our understanding of the underlying mechanisms.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis