When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
AuthorsChristopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer
Resources
This paper shows that ALiBi can make attention heads numerically blind at long distances and tests ways to restore reliable token retrieval.
Key results
Share of attention entries underflowed at token distance 2048 in bfloat16
Parameter count of the Llama decoder models used for training experiments
FineWeb-Edu tokens used for controlled pretraining
Out-of-context passkey retrieval AUC, versus 0.08 for the ALiBi baseline
What the paper found
The paper identifies a numerical failure in ALiBi, the parameter-free positional encoding used by models such as BLOOM, Falcon-RW, and MPT: its linearly increasing attention bias eventually pushes softmax exponentials below floating-point range, turning distant attention weights into exact zeros and making heads partially blind. In bfloat16, 36.6% of attention entries have underflowed by token distance 2048. The authors confirm this behavior in pretrained BLOOM 560M, Falcon-RW 7B, and MPT 7B, then isolate its effects with 148M-parameter Llama decoder models trained on 20B FineWeb-Edu tokens. Standard commonsense, question-answering, and language benchmarks change only modestly, but associative retrieval is much more sensitive: default ALiBi remains strong on needle-in-a-haystack retrieval while suffering on out-of-context passkey retrieval. Four training-time mitigations are tested—bias clamping, redesigned slopes, logarithmic distances, and soft capping. The strongest passkey result comes from combining clamping with log-scaled distances, raising out-of-context passkey AUC to 0.79 versus 0.08 for the ALiBi baseline. However, no mitigation dominates every retrieval task, and several reduce needle-in-a-haystack performance. The authors recommend clamping or explicit sliding-window attention when a bounded receptive field is intended, and view log-scaled distances as the most consistent option for passkey extrapolation.
Original abstract
We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.