NTH

Rethinking the Role of Efficient Attention in Hybrid Architectures

AuthorsZiqing Qiao, Yinuo Xu, Chaojun Xiao, Zhou Su, Zihan Zhou, Yingfa Chen, Xiaoyue Xu, Xu Han, Zhiyuan Liu

June 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that in hybrid language models, efficient attention mostly shapes how quickly long-context skills emerge, while full attention does the heavy lifting for retrieval.

Key results

41.32
S5 ShortAvg

Short-context average for SWA-128-NoPE at S5, matching the paper's 19-benchmark ShortAvg

19.46
S5 LongBench

Long-context benchmark average for SWA-128-NoPE at S5 on 21 LongBench tasks

46.98
S5 RULER 32K

Long-context RULER average for SWA-128-NoPE after 32K extension

18.30
S5 LongBench baseline

SWA-128 baseline LongBench average at S5 for comparison

41.86
S5 RULER 32K baseline

SWA-128 baseline RULER average at S5 after 32K extension

What the paper found

This paper from Tsinghua University and OpenBMB argues that efficient attention in hybrid language models is not the direct source of long-context capability; instead, it acts as an optimization prior that changes how quickly full-attention layers learn retrieval. Across seven architectures, scaling-law fits show validation Loss is nearly unchanged by the efficient-attention choice, but log(LongPPL) diverges strongly early in training and then converges as training budget increases, with SWA-2048 lagging most. Mechanistically, receptive-field restriction and Needle-in-a-Haystack probing show long-range information is carried mainly by full attention, not by sliding-window or recurrent mixers such as Lightning Attention, Mamba-2, and Gated DeltaNet. The paper names the resulting effect Large-Window Laziness: large SWA windows reduce the gradient signal that pushes full attention to form retrieval heads, delaying long-range learning. A simple design fix, applying NoPE only to full-attention layers in a small-window SWA hybrid, substantially improves long-context performance while keeping short-context quality nearly flat: at S5, ShortAvg stays at 41.32 while LongBench rises to 19.46 and RULER at 32K reaches 46.98, compared with 41.31, 18.30, and 41.86 for the SWA-128 baseline.

Original abstract

Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention modules, such as sliding-window attention (SWA) and recurrent sequence mixers. However, how these efficient modules shape model capabilities remains poorly understood. To address this gap, we conduct a systematic analysis across hybrid architectures from three perspectives: scaling behavior, mechanism analysis, and architecture design. First, from a scaling perspective, we find that efficient-attention design primarily affects how fast long-context capability emerges, while different hybrids eventually converge to comparable long-context performance under sufficient training. Second, mechanistically, we show that long-range retrieval is mainly carried by full attention, whereas efficient attention shapes its optimization trajectory. This explains a counter-intuitive phenomenon we call Large-Window Laziness: larger SWA windows can delay the formation of retrieval heads in full-attention layers. Third, guided by this mechanism, we show that applying NoPE to only the full-attention layers of a small-window SWA hybrid substantially improves long-context performance with negligible impact on short-context performance.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis