NTH

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

AuthorsTommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin, Imanol Schlag

July 10, 2026 3 min read
Watch on YouTube
The one-line take

This paper compares several modern linear-attention designs under one framework and finds that a simple cross-layer value-routing trick can modestly improve performance.

Key results

2.273
best final loss

Kimi Delta Attention with Muon in a hybrid stack on the 350M-parameter, 15B-token sweep

100.0%
fastest relative speed

Pure Gated DeltaNet with AdamW, normalized to the fastest run in the 350M, 15B-token sweep

2.433
pure Gated DeltaNet final loss

Fastest normalized training configuration in the 350M, 15B-token sweep

3.37
32k softmax iteration time

Seconds per iteration at 32k tokens

1.56
32k Gated DeltaNet hybrid iteration time

Seconds per iteration at 32k tokens

0.96
32k pure Gated DeltaNet iteration time

Seconds per iteration at 32k tokens

What the paper found

This ETH Zurich study unifies softmax attention, DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2 in a recurrent-memory formalism that exposes a clear trade-off between expressivity, memory decay, and write control, then tests that design space on 350M-parameter models trained for 15B tokens on FineWeb-Edu with the LLaMA2 tokenizer. In the main sweep, Kimi Delta Attention with Muon in a hybrid stack achieves the lowest final validation loss at 2.273, while pure Gated DeltaNet with AdamW is the fastest normalized training configuration at 100.0% relative speed but ends at 2.433 loss. Muon consistently improves matched loss across all families, hybrid stacks usually lower loss but reduce throughput, and long-context scaling strongly favors pure recurrent stacks: at 32k tokens, iteration time is 3.37 seconds for softmax, 1.56 seconds for Gated DeltaNet hybrid, and 0.96 seconds for pure Gated DeltaNet. The paper’s routing contribution is more novel than the base benchmark sweep: a DeltaNet-inspired Cross-Layer Error Residuals path does not help, but Cross-Layer Value Routing, which injects the write value into the shared residual stream through a zero-initialized projection, yields small but consistent validation-loss gains, including a 0.0103 drop for Gated DeltaNet and a 0.0119 drop for DeltaNet at 350M/1B, with smaller improvements persisting at 15B tokens and in a 1.3B/40B Gated DeltaNet run.

Original abstract

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity. Our experiments center on 350M-parameter models trained for 15B tokens, and include optimizer and learning-rate comparisons, hybrid-versus-pure stack comparisons, sequence-length runtime measurements, larger DeltaNet runs at 1.3B and 3B parameters, and a small set of downstream evaluations. The reported speed results measure training throughput and iteration time; we do not provide an empirical inference-speed benchmark. Within the reported 350M-parameter, 15B-token sweep, Kimi Delta Attention with Muon reaches the lowest final validation loss, a pure Gated DeltaNet stack trained with AdamW has the highest normalized training throughput, hybrid stacks generally improve loss at a throughput cost, and Muon consistently lowers final validation loss relative to AdamW in the matched architecture settings we evaluate. We introduce and evaluate lightweight cross-layer routing mechanisms for DeltaNet-style memories. The most natural DeltaNet-inspired formulation, forwarding a lower layer's delta-rule write error into the next layer's value target, does not improve over matched baselines. Routing into the aligned hidden stream and forwarding the write value instead yields a modest improvement in the matched runs we report: Cross-Layer Value Routing (CLVR) lowers final validation loss for both DeltaNet and Gated DeltaNet.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis