NTH

Full-bandwidth transformer

AuthorsXi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford

August 14, 2026 2 min read
Watch on YouTube
The one-line take

A transformer that feeds its hidden thoughts back into the next decoding step appears to gain better performance and shorter reasoning with little extra cost.

Key results

1B
Model size

Experiments train 1B-parameter full-bandwidth transformers.

400B
Training scale

The largest full-bandwidth models are trained on 400B tokens.

1%
Per-token decoding overhead

Latent feedback adds under 1% per-token decoding cost.

3%
Three-pass stabilization mixture

Adding 3% three-pass batches stabilizes long-horizon feedback.

0.37
MATH-500 Pass@1

Soft decoding raises the 200B-token model from 0.27 to 0.37.

67.93
GSM8K Pass@1

Instruction-tuned GSM8K improves from 64.52 to 67.93.

What the paper found

The Full-bandwidth transformer widens the narrow vertical feedback channel in autoregressive models: instead of returning only the sampled token, it fuses that token’s embedding with the previous top-layer hidden state using a gated linear unit, then feeds the result into the next decoding step. This lets non-verbalized plans, uncertainty, and intermediate computations receive a fresh traversal through the full stack while preserving the standard architecture, KV cache, language-modeling objective, and Microsoft-compatible serving workflows. The added decoding cost is under 1% per token, requiring only two projection multiplications. To train the recurrence without sacrificing parallel teacher forcing, the method uses scheduled multi-pass training, prefix mixin, RMS normalization, tied input-output embeddings, and hidden-state jitter; notably, adding 3% three-pass batches stabilizes feedback beyond the training horizon. Experiments use 1B-parameter models trained on up to 400B tokens with the Phi-4 data mixture, and optional fused prefilling roughly doubles data efficiency. On the 200B-token model, Soft decoding raises MATH-500 Pass@1 from 0.27 to 0.37, while instruction-tuned GSM8K improves from 64.52 to 67.93 and HumanEval reaches 45.92 Pass@3. Synthetic state-tracking probes show that one recurrent step makes information nearly immediately accessible, reaching 99.6% completion accuracy and 100% delayed-memory accuracy at the input layer. The approach outperforms matched standard baselines and compares favorably with models such as Llama3.2 1B and Qwen3 1.7B, while often producing shorter reasoning traces than conventional chain-of-thought.

Original abstract

Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce the \emph{full-bandwidth transformer}, which widens this channel with \emph{latent feedback}: at each decoding step, the previous top-layer hidden state is fused with the sampled token embedding through a gated linear unit and fed back as the next input. Latent feedback lets non-verbalized computation re-enter the stack with a renewed depth budget, while preserving the standard transformer architecture, KV cache, and language-modeling objective. To train full-bandwidth transformers without losing parallel teacher forcing, we use a scheduled multi-pass objective that introduces latent feedback late in pretraining and mixes a small fraction of deeper feedback passes for stability. We train 1B-parameter full-bandwidth transformers up to 400B tokens and find that latent feedback improves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance. With negligible per-token decoding overhead, full-bandwidth transformers match or approach standard transformers trained with roughly $1.5\times$ more tokens, and manage to produce shorter reasoning traces at equal or better accuracy.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis