Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
AuthorsTommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin, Imanol Schlag
Resources
This paper compares several modern linear-attention designs under one framework and finds that a simple cross-layer value-routing trick can modestly improve performance.
Key results
Kimi Delta Attention with Muon in a hybrid stack on the 350M-parameter, 15B-token sweep
Pure Gated DeltaNet with AdamW, normalized to the fastest run in the 350M, 15B-token sweep
Fastest normalized training configuration in the 350M, 15B-token sweep
Seconds per iteration at 32k tokens
Seconds per iteration at 32k tokens
Seconds per iteration at 32k tokens
What the paper found
This ETH Zurich study unifies softmax attention, DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2 in a recurrent-memory formalism that exposes a clear trade-off between expressivity, memory decay, and write control, then tests that design space on 350M-parameter models trained for 15B tokens on FineWeb-Edu with the LLaMA2 tokenizer. In the main sweep, Kimi Delta Attention with Muon in a hybrid stack achieves the lowest final validation loss at 2.273, while pure Gated DeltaNet with AdamW is the fastest normalized training configuration at 100.0% relative speed but ends at 2.433 loss. Muon consistently improves matched loss across all families, hybrid stacks usually lower loss but reduce throughput, and long-context scaling strongly favors pure recurrent stacks: at 32k tokens, iteration time is 3.37 seconds for softmax, 1.56 seconds for Gated DeltaNet hybrid, and 0.96 seconds for pure Gated DeltaNet. The paper’s routing contribution is more novel than the base benchmark sweep: a DeltaNet-inspired Cross-Layer Error Residuals path does not help, but Cross-Layer Value Routing, which injects the write value into the shared residual stream through a zero-initialized projection, yields small but consistent validation-loss gains, including a 0.0103 drop for Gated DeltaNet and a 0.0119 drop for DeltaNet at 350M/1B, with smaller improvements persisting at 15B tokens and in a 1.3B/40B Gated DeltaNet run.
Original abstract
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity. Our experiments center on 350M-parameter models trained for 15B tokens, and include optimizer and learning-rate comparisons, hybrid-versus-pure stack comparisons, sequence-length runtime measurements, larger DeltaNet runs at 1.3B and 3B parameters, and a small set of downstream evaluations. The reported speed results measure training throughput and iteration time; we do not provide an empirical inference-speed benchmark. Within the reported 350M-parameter, 15B-token sweep, Kimi Delta Attention with Muon reaches the lowest final validation loss, a pure Gated DeltaNet stack trained with AdamW has the highest normalized training throughput, hybrid stacks generally improve loss at a throughput cost, and Muon consistently lowers final validation loss relative to AdamW in the matched architecture settings we evaluate. We introduce and evaluate lightweight cross-layer routing mechanisms for DeltaNet-style memories. The most natural DeltaNet-inspired formulation, forwarding a lower layer's delta-rule write error into the next layer's value target, does not improve over matched baselines. Routing into the aligned hidden stream and forwarding the write value instead yields a modest improvement in the matched runs we report: Cross-Layer Value Routing (CLVR) lowers final validation loss for both DeltaNet and Gated DeltaNet.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.