HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
AuthorsZhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
AffiliationsMonash University · Zhejiang University
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
Key results
Maximum improvement over native Gated DeltaNet across Qwen3.5 model scales.
Maximum macro-average gain over native Gated DeltaNet on Qwen3.5.
Model used for controlled long-context generalization experiments.
SlimPajama training budget for the from-scratch comparison.
HLA improvement over Gated DeltaNet when trained at 4K and evaluated at 32K context.
Reduction achieved with the default 256-token chunk size versus the matched KV cache.
What the paper found
HLA, or Hybrid Linear Attention, extends Gated DeltaNet by making historical memory access query-dependent rather than fixed. It summarizes each completed chunk as an exact affine state transition, pools compact representatives with local self-attention, and uses sigmoid routing gates to interpolate each transition with the identity, controlling both the chunk’s additive memory and its transformation of earlier states. Effective-support regularization encourages sparse routing, while inference removes gates below 0.1. Unlike recurrent approaches such as Mamba, which compress history into a largely query-independent trajectory, HLA dynamically recomposes chunk summaries for each query without reverting to full softmax attention. On pretrained Qwen3.5 models from 0.8B to 9B, HLA improves LongBench-V2 by up to 5.57 percentage points and RULER by up to 3.974 points over native Gated DeltaNet, consistently outperforming fixed chunk mixing through MHLA. In a matched from-scratch experiment, a 1.3B model trained on SlimPajama for 100B tokens with a 4K context gains 4.22 points on RULER at 32K context, demonstrating extrapolation beyond training length. With a 256-token chunk size, HLA also cuts affine historical-state storage by 2x relative to the matched KV cache, although query routing adds moderate decoding overhead.
Original abstract
Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: https://caesarhhh.github.io/hla/
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.
Block Sparse Attention with Log-Linear Complexity
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu
PISA makes long-context attention more scalable by hierarchically narrowing relevant key blocks instead of comparing every query with every block.