NTH

HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization

AuthorsZhentao Tan, Wei Chen, Jingyi Shen, Yao Liu, Xu Shen, Yue Wu, Jieping Ye

July 4, 2026 3 min read
Watch on YouTube
The one-line take

HydraHead rethinks hybrid attention by assigning full attention only to the most retrieval-critical heads, improving long-context performance while keeping compute efficient.

Key results

15B
training tokens

HydraHead scaling run used only 15B tokens

7:1
LA-to-FA ratio

HydraHead matched the 3:1 layer-wise hybrid at this higher linear-to-full attention ratio

94.53
RULER Single 256K

HydraHead score on RULER Single-Key at 256K context

52.70
RULER Multi-Key 256K

HydraHead score on RULER Multi-Key at 256K context

50.62
general reasoning average

HydraHead average on MMLU, BBH, MBPP, and GSM8K

What the paper found

HydraHead from Alibaba Group reframes long-context hybrid Transformers by showing that attention heads, not layers, are the right granularity for mixing Full Attention and Linear Attention. Using mechanistic interpretability on Qwen3-1.7B, the authors find that retrieval-critical behavior is sparse, causally localized, and scattered across layers rather than aligned to layer boundaries, so they assign Full Attention only to the heads with the highest causal importance and route the rest through Gated DeltaNet linear attention. To reconcile the mismatch between softmax and linear-attention outputs, HydraHead adds head-wise scale-normalized fusion plus a three-stage transfer pipeline: parameter reuse and layer alignment, global logit distillation, and long-context fine-tuning. Under a unified setup, the model outperforms layer-wise, token-wise, and other head-wise hybrids, matching a 3:1 layer-wise hybrid’s long-context performance even at a 7:1 LA-to-FA ratio. In the optimized scaling run, trained on 15B tokens, HydraHead reaches 94.53% on RULER Single-Key at 256K and 52.70% on RULER Multi-Key at 256K, while remaining competitive on general reasoning with a 50.62 average across MMLU, BBH, MBPP, and GSM8K. The paper’s key technical claim is that interpretability-guided head selection enables aggressive Full Attention compression without collapsing retrieval, yielding a more balanced long-context architecture than current layer-wise hybrid designs.

Original abstract

The quadratic complexity of attention poses a critical bottleneck for long-context processing, spurring interest in hybrid attention designs. Most open-source hybrid models adopt a layer-wise strategy. Yet, prior work has noted the inherent difficulty of integrating Linear Attention (LA) with Full Attention (FA), suggesting that the design space of attention hybridization remains underexplored. To probe this space, we conduct interpretability analysis and observe that layers exhibit block-wise functional similarity, while individual heads within the same layer display distinct functional specialization despite sharing input features. This head-level heterogeneity suggests that the head dimension provides a natural and principled granularity for fusing heterogeneous attention signals. Building on this insight, we introduce HydraHead, a novel architecture that hybridizes FA and LA along the head axis. HydraHead features two key innovations: (1) an interpretability-driven selection strategy that identifies retrieval-critical heads and preserves FA only for them, and (2) a scale-normalized fusion module that reconciles the distributional gap between FA and LA head outputs. By leveraging a three-stage transfer pipeline with parameter reuse and distillation, we achieve high-performance hybrid models with minimal training overhead. Under a unified training setup, HydraHead outperforms other hybrid designs in long-context tasks while maintaining strong general reasoning. With interpretability-driven head selection, it matches a 3:1 layer-wise hybrid's long-context performance at a 7:1 LA-to-FA ratio. Crucially, trained on only 15B tokens, HydraHead achieves over 69% improvement over the baseline at 512K context length, approaching Qwen3.5, a leading model of comparable size with a native context length of 256K. This highlights the significant scaling potential of head-level hybridization.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis