NTH

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

AuthorsChristopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer

August 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that ALiBi can make attention heads numerically blind at long distances and tests ways to restore reliable token retrieval.

Key results

36.6%
Underflow fraction

Share of attention entries underflowed at token distance 2048 in bfloat16

148M
Controlled model size

Parameter count of the Llama decoder models used for training experiments

20B
Training corpus scale

FineWeb-Edu tokens used for controlled pretraining

0.79
Clamping plus log passkey AUC

Out-of-context passkey retrieval AUC, versus 0.08 for the ALiBi baseline

What the paper found

The paper identifies a numerical failure in ALiBi, the parameter-free positional encoding used by models such as BLOOM, Falcon-RW, and MPT: its linearly increasing attention bias eventually pushes softmax exponentials below floating-point range, turning distant attention weights into exact zeros and making heads partially blind. In bfloat16, 36.6% of attention entries have underflowed by token distance 2048. The authors confirm this behavior in pretrained BLOOM 560M, Falcon-RW 7B, and MPT 7B, then isolate its effects with 148M-parameter Llama decoder models trained on 20B FineWeb-Edu tokens. Standard commonsense, question-answering, and language benchmarks change only modestly, but associative retrieval is much more sensitive: default ALiBi remains strong on needle-in-a-haystack retrieval while suffering on out-of-context passkey retrieval. Four training-time mitigations are tested—bias clamping, redesigned slopes, logarithmic distances, and soft capping. The strongest passkey result comes from combining clamping with log-scaled distances, raising out-of-context passkey AUC to 0.79 versus 0.08 for the ALiBi baseline. However, no mitigation dominates every retrieval task, and several reduce needle-in-a-haystack performance. The authors recommend clamping or explicit sliding-window attention when a bounded receptive field is intended, and view log-scaled distances as the most consistent option for passkey extrapolation.

Original abstract

We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis