NTH

A Controlled Study of Attention-Only Transformers

AuthorsHenry Ndubuaku, Karen Mosoyan, Jakub Mroz, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H Lee

July 25, 2026 3 min read
Watch on YouTube
The one-line take

A large controlled study finds that attention alone can nearly reproduce transformer performance, though feed-forward layers still provide an advantage for recalling knowledge stored in the weights.

Key results

105B
Training token budget

Largest pretraining budget used for the main SAN and FFN comparisons.

0.470
Matched-depth FFN removal cost

Increase in validation loss, measured in nats, when the FFN is deleted without reallocating capacity.

0.263
Matched-FLOP loss gap

Validation-loss advantage of the standard FFN transformer at equal training compute.

0.0055
Matched-parameter loss gap

Clean seed-pair gap after reallocating the FFN parameter budget into attention depth.

0.27%
Relative matched-parameter gap

Matched-parameter loss difference as a percentage of total loss.

0.0398
fineweb-edu gap

Measured validation-loss gap on the knowledge-dense fineweb-edu control pair.

What the paper found

Researchers at Cactus Compute conduct what they describe as the first controlled necessity test of the transformer feed-forward network, comparing standard decoder transformers with attention-only Simple Attention Networks, or SANs. Across 2 to 48 layers, 6M to 87M parameters, and up to 105B tokens from the reasoning-dense SYNTH corpus, they separately match parameter count, training FLOPs, and depth while tuning learning rates for every arm. Deleting the FFN in place costs 0.470 nats at matched depth and 0.263 nats at matched FLOPs, because attention’s quadratic computation leaves fewer parameters for storage. Reallocating that budget into attention depth nearly eliminates the difference: at matched parameter count, the clean validation gap is 0.0055 nats, or 0.27% of loss. The residual deficit concentrates on low-context prediction, where the model cannot route useful information from the prompt; SANs instead match or outperform FFN models on context-grounded tasks such as SciQ. A pre-registered test on the knowledge-dense fineweb-edu dataset measured a 0.0398-nat SAN deficit, while the distribution-sensitive LAMBADA results favored FFN models for out-of-distribution recall. Mechanistically, QK-normalization is essential for training deep SANs, whereas residual gating is performance-neutral and sandwich normalization improves loss. Weight spectra show Q/K routing matrices crystallizing early, while content-carrying projections accumulate rank throughout training, indicating that the FFN mainly supplies parametric memory and that attention can recover most of its function when given equivalent parameter capacity.

Original abstract

Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder transformers (Simple Attention Networks, SANs) against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters. Deleting feed-forward layers in place is costly: the standard transformer leads by 0.47 nats at matched depth and 0.26 nats at matched FLOPs. Reallocating the freed budget into attention depth closes the gap: at matched parameters the difference is 0.006 nats (0.27 percent of loss), reproducible to one part in ten thousand across seed pairs, shrinking across 5B, 30B, and 105B budgets, and holding near 0.02 nats across a 29x size range. Three measurements localize the remaining gap to parametric recall: attention-only models are better on context-grounded answers and worse where knowledge must come from weights. Weight spectra show why: routing matrices (Q/K) crystallize early, content matrices accumulate rank slowly, and removing feed-forward layers relocates this accumulation to the attention output projection. QK-normalization, not feed-forward layers or residual gating, keeps 48-layer attention-only stacks trainable. The deficit concentrates on low-context query prediction and localizes there entirely by the largest budget. A pre-registered test confirms the account: it predicts a 0.02 to 0.05 nat gap on knowledge-dense web text; a matched pair trained on fineweb-edu measures 0.040. Within the tested regime, attention does the rest.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis