Resources
Maglev teaches a lightweight recurrent Transformer to preserve long-range context in a compact memory, enabling faster long-context generation without full attention at inference.
Key results
Size of the deployed nanochat d20 Maglev decoder
Pretraining budget used in the main experiments
Fixed token and recurrent-memory window used by decoder P
Best score from separate-parameter Maglev with λ = 1
Best mean accuracy across the reported pretraining benchmarks
What the paper found
Maglev is a recurrent Transformer that adds persistent, token-level memory to sliding-window attention without sacrificing parallel pretraining. A causal prefiller, Q, uses broader attention to construct memory targets in parallel, while decoder P uses a fixed 512-token window and gated recurrent key/value injection to reproduce those memories and predict the next token. A consistency loss aligns P’s memory with Q’s target; at inference, Q is discarded, and P feeds back its own memories with a cache that remains bounded regardless of sequence length. In the nanochat d20 setup, the deployed decoder has 435M parameters, training uses 43.52B tokens, and the maximum context is 2048 tokens. Against a matched sliding-window Transformer, the best separate-parameter Maglev configuration, with consistency weight λ = 1, reduces FineWeb-Edu validation BPB from 0.7413 to 0.7251 and raises average accuracy across downstream benchmarks from 54.1 to 56.4, while also outperforming latent recurrent Transformer baselines. Parameter sharing between Q and P is feasible: the shared model with λ = 0.1 reaches 0.7295 FineWeb-Edu BPB and 56.2 average downstream accuracy, retaining most of the gain with lower parameter memory. The central contribution is a nonlinear recurrent memory interface trained through two sequence-parallel passes, combining Transformer-style representation capacity with sliding-window inference costs.
Original abstract
We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. \ours{} consists of two coupled models: a prefiller $Q$, which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for $Q$, as this yields stronger performance. The essential requirement is that $Q$ be more expressive than $P$, with access to the full history.} to produce memory targets $m'_t$, and a decoder $P$, which uses only sliding-window attention and recurrent K/V injection to produce decoder memories $m_t$ for next-token prediction. We train \ours{} with a memory consistency loss that aligns $m_t$ with $m'_t$, allowing inference to use $P$ alone. Empirically, \ours{} improves validation loss and downstream pretraining benchmarks over sliding-window and latent recurrent transformer baselines. Moreover, sharing parameters between $P$ and $Q$ reduces parameter memory while preserving most of the gains.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.