LT2: Linear-Time Looped Transformers
AuthorsChunyuan Deng, Yizhe Zhang, Rui-Jie Zhu, Yuanyuan Xu, Jiarui Liu, T. S. Eugene Ng, Hanjie Chen
Resources
This paper makes looped transformers much cheaper by swapping in linear or sparse attention, showing you can keep quality while dramatically cutting compute.
Key results
LT2 models are pretrained from scratch at two scales on FineWeb-Edu with the same 100B-token budget.
All looped variants use 4 loops over the shared physical layer stack during pretraining and evaluation.
At 1.3B, the linear-sparse hybrid matches the standard looped transformer’s quality while avoiding quadratic attention.
LT2-hybrid (GDN+DSA) reaches 125 tokens/s decode throughput at 8k context, compared with 22 tokens/s for the baseline.
LT2-hybrid (Full+GDN) improves average zero-shot accuracy by 2.1 points over the standard Looped Transformer.
The converted Ouro-Hybrid-1.4B model is obtained with about 1B continuation tokens from a pretrained looped transformer.
What the paper found
LT2, or Linear-Time Looped Transformers, replaces the quadratic softmax attention bottleneck in looped transformers with subquadratic token mixers while preserving weight-sharing across iterations. The paper studies LT2-linear, built from linear-attention primitives such as Gated Delta Networks (GDN), KDA, and Mamba2-style recurrences, and LT2-sparse, built from sparse attention such as Native Sparse Attention (NSA) and DeepSeek’s DSA. The key novelty is that looping is not just compatible with efficient mixers; it amplifies them. In the linear case, repeated application upgrades a rank-1 memory update into an effective rank-T update, and in the sparse case, T loops expand a sliding window’s receptive field to roughly T×w tokens. On FineWeb-Edu pretraining at 0.6B and 1.3B parameters with T=4, LT2 variants stay within about 1 average zero-shot point of the full-attention loop, while LT2-hybrid (GDN+DSA) matches the standard Looped Transformer at 1.3B perplexity 9.72 versus 9.87 and reaches 125 tokens/s decoding at 8k context, about 5.7× faster than the 22 tokens/s baseline. The strongest model, LT2-hybrid (Full+GDN), improves average zero-shot accuracy by 2.1 points to 61.39 and achieves 2.7× faster decoding, establishing a new quality-efficiency Pareto frontier. A converted 1.4B model, Ouro-Hybrid-1.4B, obtained with about 1B continuation tokens, outperforms industry-level 1B models and approaches 4B-class models on benchmarks including ARC-C, HellaSwag, WinoGrande, MMLU, GSM8K, and RULER. The paper also shows that GDN stabilizes training better than RetNet or DeltaNet, and that a simple SDPA output gate reduces attention-sink compounding across loops.
Original abstract
Looped Transformers (LT) have emerged as a powerful architecture by iterating their layers multiple times before decoding the final token. However, pairing them with full attention retains quadratic complexity, making them computationally expensive and slow. We introduce LT2 (Linear-Time Looped Transformers), a family of looped architectures that replace quadratic softmax attention with subquadratic, linear-time attention. We study two variants: LT2-linear with linear attention and LT2-sparse with sparse attention. We find that looping uniquely synergizes with these variants: it enables iterative memory refinement in linear attention and progressively expands the effective receptive field in sparse attention. We formalize these benefits theoretically and demonstrate consistent empirical gains across controlled recall, state-tracking, and language modeling tasks. We then explore LT2-hybrid, which combines different attention variants in a looped setting. Two variants are especially promising: LT2-hybrid (GDN+DSA), which interleaves linear and sparse attention to maximize efficiency and matches the standard looped transformer's quality at fully linear-time cost; and LT2-hybrid (Full+GDN), which interleaves GDN with a small fraction of full attention layers to maximize quality, surpassing the standard looped transformer in both performance and efficiency. We also show how to convert a pre-trained LT into an LT2-hybrid model. With about 1B tokens of training, our converted model, Ouro-hybrid-1.4B, outperforms industry-level 1B models and is competitive with industry-level 4B models while retaining the speed benefits of linear-time attention. Together, these results show a clear path toward making looped transformers more scalable and advancing efficient, capable small language models.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.