A Controlled Study of Attention-Only Transformers
AuthorsHenry Ndubuaku, Karen Mosoyan, Jakub Mroz, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H Lee
Resources
A large controlled study finds that attention alone can nearly reproduce transformer performance, though feed-forward layers still provide an advantage for recalling knowledge stored in the weights.
Key results
Largest pretraining budget used for the main SAN and FFN comparisons.
Increase in validation loss, measured in nats, when the FFN is deleted without reallocating capacity.
Validation-loss advantage of the standard FFN transformer at equal training compute.
Clean seed-pair gap after reallocating the FFN parameter budget into attention depth.
Matched-parameter loss difference as a percentage of total loss.
Measured validation-loss gap on the knowledge-dense fineweb-edu control pair.
What the paper found
Researchers at Cactus Compute conduct what they describe as the first controlled necessity test of the transformer feed-forward network, comparing standard decoder transformers with attention-only Simple Attention Networks, or SANs. Across 2 to 48 layers, 6M to 87M parameters, and up to 105B tokens from the reasoning-dense SYNTH corpus, they separately match parameter count, training FLOPs, and depth while tuning learning rates for every arm. Deleting the FFN in place costs 0.470 nats at matched depth and 0.263 nats at matched FLOPs, because attention’s quadratic computation leaves fewer parameters for storage. Reallocating that budget into attention depth nearly eliminates the difference: at matched parameter count, the clean validation gap is 0.0055 nats, or 0.27% of loss. The residual deficit concentrates on low-context prediction, where the model cannot route useful information from the prompt; SANs instead match or outperform FFN models on context-grounded tasks such as SciQ. A pre-registered test on the knowledge-dense fineweb-edu dataset measured a 0.0398-nat SAN deficit, while the distribution-sensitive LAMBADA results favored FFN models for out-of-distribution recall. Mechanistically, QK-normalization is essential for training deep SANs, whereas residual gating is performance-neutral and sandwich normalization improves loss. Weight spectra show Q/K routing matrices crystallizing early, while content-carrying projections accumulate rank throughout training, indicating that the FFN mainly supplies parametric memory and that attention can recover most of its function when given equivalent parameter capacity.
Original abstract
Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder transformers (Simple Attention Networks, SANs) against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters. Deleting feed-forward layers in place is costly: the standard transformer leads by 0.47 nats at matched depth and 0.26 nats at matched FLOPs. Reallocating the freed budget into attention depth closes the gap: at matched parameters the difference is 0.006 nats (0.27 percent of loss), reproducible to one part in ten thousand across seed pairs, shrinking across 5B, 30B, and 105B budgets, and holding near 0.02 nats across a 29x size range. Three measurements localize the remaining gap to parametric recall: attention-only models are better on context-grounded answers and worse where knowledge must come from weights. Weight spectra show why: routing matrices (Q/K) crystallize early, content matrices accumulate rank slowly, and removing feed-forward layers relocates this accumulation to the attention output projection. QK-normalization, not feed-forward layers or residual gating, keeps 48-layer attention-only stacks trainable. The deficit concentrates on low-context query prediction and localizes there entirely by the largest budget. A pre-registered test confirms the account: it predicts a 0.02 to 0.05 nat gap on knowledge-dense web text; a matched pair trained on fineweb-edu measures 0.040. Within the tested regime, attention does the rest.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.