Dynamic Linear Attention
AuthorsXin Wang, Hui Shen, Boyuan Zheng, Xueshen Liu, Minkyoung Cho, Zhongwei Wan, Zesen Zhao, Zhuoqing Mao, Shen Yan, Mi Zhang
Resources
This paper makes linear attention smarter by dynamically deciding how to merge memory states, aiming to keep long-context models efficient without losing important information.
Key results
Academic-scale pretraining data used for DLA variants
Training context length
Fixed maximum number of memory states in DLA
Average relative improvement of DLA over Log-Linear Attention on Mamba-2 across eight commonsense benchmarks
Mamba-2 with DLA on the long-context retrieval benchmark
What the paper found
Dynamic Linear Attention, or DLA, is a new multi-state linear-attention framework from The Ohio State University, the University of Michigan, and ByteDance Seed that targets long-context LLM scaling without the rigid merge schedules used by Log-Linear Attention. The core idea is to replace fixed block construction with information-aware dynamic state merging: DLA computes a State Information Score from the Frobenius-norm drift between the current token state and the most recent memory state, then starts a new state only when that score crosses a threshold; otherwise it keeps accumulating into the existing state. It also adds capacity-bounded memory modeling, maintaining a fixed cache of 30 states and greedily merging adjacent low-information states when full, which keeps inference cost predictable while preserving temporal order. The method was pretrained on 50B tokens with 16K sequence length on Mamba-2-780M and Gated DeltaNet-1.3B, then evaluated on 16 datasets across commonsense reasoning, in-context retrieval, and long-context benchmarks. DLA consistently beats Log-Linear Attention, including an 8% average gain on commonsense tasks for Mamba-2 and 6% for Gated DeltaNet, and it reaches 63.1 on RULER MV-NIAH versus 28.3 for Log-Linear on Mamba-2. Efficiency experiments also show higher throughput and lower runtime memory than Log-Linear Attention, while ablations confirm that both the dynamic boundary rule and the bounded cache contribute meaningfully to the final gains.
Original abstract
The scalability of Large Language Models (LLMs) to long contexts is fundamentally constrained by the quadratic complexity of standard attention, motivating the adoption of linear attention mechanisms with sub-quadratic cost. To improve representation capacity under long contexts, recent approaches organize memory in a multi-state manner. However, existing multi-state linear attention methods rely on fixed state merging policies that cannot adapt to dynamically varying token importance, irreversibly obscuring critical tokens and causing severe error accumulation over long sequences. To address this limitation, we propose DLA, a dynamic memory modeling framework for multi-state linear attention. DLA introduces (i) Information-Aware Dynamic State Merging, which adaptively determines state boundaries based on token-level information variation, preserving high-resolution representations around semantic transitions while aggressively summarizing stable regions, and (ii) Capacity-Bounded Memory Modeling, which maintains a fixed-size, chronologically ordered state cache by selectively merging adjacent low-information states to control memory growth with minimal information loss. We pre-train DLA on two different linear attention models and evaluate on 16 datasets across three categories. Experimental results demonstrate the superiority of DLA over state-of-the-art.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.