NTH

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

AuthorsZhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang

AffiliationsMonash University · Zhejiang University

October 8, 2026 2 min read
Watch on YouTube
The one-line take

HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.

Key results

5.57
LongBench-V2 improvement

Maximum improvement over native Gated DeltaNet across Qwen3.5 model scales.

3.974
RULER improvement

Maximum macro-average gain over native Gated DeltaNet on Qwen3.5.

1.3B
From-scratch model size

Model used for controlled long-context generalization experiments.

100B
Training token budget

SlimPajama training budget for the from-scratch comparison.

4.22
RULER gain at 32K

HLA improvement over Gated DeltaNet when trained at 4K and evaluated at 32K context.

2x
Historical-state storage reduction

Reduction achieved with the default 256-token chunk size versus the matched KV cache.

What the paper found

HLA, or Hybrid Linear Attention, extends Gated DeltaNet by making historical memory access query-dependent rather than fixed. It summarizes each completed chunk as an exact affine state transition, pools compact representatives with local self-attention, and uses sigmoid routing gates to interpolate each transition with the identity, controlling both the chunk’s additive memory and its transformation of earlier states. Effective-support regularization encourages sparse routing, while inference removes gates below 0.1. Unlike recurrent approaches such as Mamba, which compress history into a largely query-independent trajectory, HLA dynamically recomposes chunk summaries for each query without reverting to full softmax attention. On pretrained Qwen3.5 models from 0.8B to 9B, HLA improves LongBench-V2 by up to 5.57 percentage points and RULER by up to 3.974 points over native Gated DeltaNet, consistently outperforming fixed chunk mixing through MHLA. In a matched from-scratch experiment, a 1.3B model trained on SlimPajama for 100B tokens with a 4K context gains 4.22 points on RULER at 32K context, demonstrating extrapolation beyond training length. With a 256-token chunk size, HLA also cuts affine historical-state storage by 2x relative to the matched KV cache, although query routing adds moderate decoding overhead.

Original abstract

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: https://caesarhhh.github.io/hla/

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis
03Attention

Block Sparse Attention with Log-Linear Complexity

Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu

PISA makes long-context attention more scalable by hierarchically narrowing relevant key blocks instead of comparing every query with every block.

Read analysis