Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
AuthorsZunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo
This study shows that unusual activation spikes in hybrid-attention language models follow predictable patterns tied to where full attention occurs and how activations are canceled.
Key results
Representative hybrid configurations achieve perfect localization of PAS at consensus attention sinks.
The cross-model analysis includes Qwen3.5 up to 397B total parameters.
Controlled Gated DeltaNet pretraining shows PAS after 1B training tokens.
Inter-spike retention for the 340M Gated DeltaNet model under a 3:1 hybridization.
Inter-spike retention for the 1.3B Gated DeltaNet model under the same configuration.
What the paper found
This study examines massive activations—sparse, unusually large hidden-state values associated with attention-sink tokens—in hybrid linear-attention language models. Across five linear-attention architectures, six hybridization configurations, and five domains including WikiText-103, GSM8K, CodeSearchNet, Scientific Papers, and FLORES-200, it identifies two architecture-aligned patterns: pre-attention spikes, or PAS, where activations peak immediately before full-attention layers, and inter-spike plateaus, or ISP, where those values persist through intervening linear-attention layers. At denser full-attention schedules, ISP increasingly connects successive PAS, converging toward the stable massive-activation profile of conventional Transformers. The pattern recurs in Qwen3.5, Kimi Linear, Nemotron-H, and Zamba2, spanning models from 1.2B to 397B parameters and extending from linear-attention to Mamba-based hybrids. Sink–spike alignment reaches 100.0% in representative configurations. Controlled Gated DeltaNet pretraining shows PAS emerging after 1B tokens, while ISP retention rises from 90.56% at 340M parameters to 94.47% at 1.3B. Mechanistically, fixed-coordinate analysis supports a write–sink–cancel lifecycle: a pre-attention layer writes an outlier, full attention concentrates on its token as a sink, and an opposite-signed update cancels it; ISP reflects delayed cancellation. Adding output gates to full-attention layers sharply reduces activation magnitude but does not remove the layerwise organization, whereas removing Gated DeltaNet gates has a comparatively modest effect.
Original abstract
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention architectures, six hybridization configurations, five data domains, and representative open-source hybrid models spanning 1.2B to 397B total parameters. Controlled pretraining of GDN-based hybrids at scales up to 1.3B shows that both morphologies emerge early and respond asymmetrically to output gating: full attention output gating strongly attenuates their absolute magnitudes without eliminating their layerwise organization, whereas removing GDN gates yields comparatively modest amplification. Mechanistically, our systematic-outlier analysis supports a shared lifecycle account governed by the timing of MA cancellation. PAS follows a localized write-sink-cancel process, while the extended persistence of ISP is consistent with delayed cancellation. At the full attention limit, this account recovers the stable MA morphology characteristic of full attention LLMs. Our code is available at https://github.com/StartluxLabs/Massive-Activations-HLA.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.