Rethinking the Role of Efficient Attention in Hybrid Architectures
AuthorsZiqing Qiao, Yinuo Xu, Chaojun Xiao, Zhou Su, Zihan Zhou, Yingfa Chen, Xiaoyue Xu, Xu Han, Zhiyuan Liu
Resources
This paper shows that in hybrid language models, efficient attention mostly shapes how quickly long-context skills emerge, while full attention does the heavy lifting for retrieval.
Key results
Short-context average for SWA-128-NoPE at S5, matching the paper's 19-benchmark ShortAvg
Long-context benchmark average for SWA-128-NoPE at S5 on 21 LongBench tasks
Long-context RULER average for SWA-128-NoPE after 32K extension
SWA-128 baseline LongBench average at S5 for comparison
SWA-128 baseline RULER average at S5 after 32K extension
What the paper found
This paper from Tsinghua University and OpenBMB argues that efficient attention in hybrid language models is not the direct source of long-context capability; instead, it acts as an optimization prior that changes how quickly full-attention layers learn retrieval. Across seven architectures, scaling-law fits show validation Loss is nearly unchanged by the efficient-attention choice, but log(LongPPL) diverges strongly early in training and then converges as training budget increases, with SWA-2048 lagging most. Mechanistically, receptive-field restriction and Needle-in-a-Haystack probing show long-range information is carried mainly by full attention, not by sliding-window or recurrent mixers such as Lightning Attention, Mamba-2, and Gated DeltaNet. The paper names the resulting effect Large-Window Laziness: large SWA windows reduce the gradient signal that pushes full attention to form retrieval heads, delaying long-range learning. A simple design fix, applying NoPE only to full-attention layers in a small-window SWA hybrid, substantially improves long-context performance while keeping short-context quality nearly flat: at S5, ShortAvg stays at 41.32 while LongBench rises to 19.46 and RULER at 32K reaches 46.98, compared with 41.31, 18.30, and 41.86 for the SWA-128 baseline.
Original abstract
Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention modules, such as sliding-window attention (SWA) and recurrent sequence mixers. However, how these efficient modules shape model capabilities remains poorly understood. To address this gap, we conduct a systematic analysis across hybrid architectures from three perspectives: scaling behavior, mechanism analysis, and architecture design. First, from a scaling perspective, we find that efficient-attention design primarily affects how fast long-context capability emerges, while different hybrids eventually converge to comparable long-context performance under sufficient training. Second, mechanistically, we show that long-range retrieval is mainly carried by full attention, whereas efficient attention shapes its optimization trajectory. This explains a counter-intuitive phenomenon we call Large-Window Laziness: larger SWA windows can delay the formation of retrieval heads in full-attention layers. Third, guided by this mechanism, we show that applying NoPE to only the full-attention layers of a small-window SWA hybrid substantially improves long-context performance with negligible impact on short-context performance.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.