Full-bandwidth transformer
AuthorsXi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
Resources
A transformer that feeds its hidden thoughts back into the next decoding step appears to gain better performance and shorter reasoning with little extra cost.
Key results
Experiments train 1B-parameter full-bandwidth transformers.
The largest full-bandwidth models are trained on 400B tokens.
Latent feedback adds under 1% per-token decoding cost.
Adding 3% three-pass batches stabilizes long-horizon feedback.
Soft decoding raises the 200B-token model from 0.27 to 0.37.
Instruction-tuned GSM8K improves from 64.52 to 67.93.
What the paper found
The Full-bandwidth transformer widens the narrow vertical feedback channel in autoregressive models: instead of returning only the sampled token, it fuses that token’s embedding with the previous top-layer hidden state using a gated linear unit, then feeds the result into the next decoding step. This lets non-verbalized plans, uncertainty, and intermediate computations receive a fresh traversal through the full stack while preserving the standard architecture, KV cache, language-modeling objective, and Microsoft-compatible serving workflows. The added decoding cost is under 1% per token, requiring only two projection multiplications. To train the recurrence without sacrificing parallel teacher forcing, the method uses scheduled multi-pass training, prefix mixin, RMS normalization, tied input-output embeddings, and hidden-state jitter; notably, adding 3% three-pass batches stabilizes feedback beyond the training horizon. Experiments use 1B-parameter models trained on up to 400B tokens with the Phi-4 data mixture, and optional fused prefilling roughly doubles data efficiency. On the 200B-token model, Soft decoding raises MATH-500 Pass@1 from 0.27 to 0.37, while instruction-tuned GSM8K improves from 64.52 to 67.93 and HumanEval reaches 45.92 Pass@3. Synthetic state-tracking probes show that one recurrent step makes information nearly immediately accessible, reaching 99.6% completion accuracy and 100% delayed-memory accuracy at the input layer. The approach outperforms matched standard baselines and compares favorably with models such as Llama3.2 1B and Qwen3 1.7B, while often producing shorter reasoning traces than conventional chain-of-thought.
Original abstract
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce the \emph{full-bandwidth transformer}, which widens this channel with \emph{latent feedback}: at each decoding step, the previous top-layer hidden state is fused with the sampled token embedding through a gated linear unit and fed back as the next input. Latent feedback lets non-verbalized computation re-enter the stack with a renewed depth budget, while preserving the standard transformer architecture, KV cache, and language-modeling objective. To train full-bandwidth transformers without losing parallel teacher forcing, we use a scheduled multi-pass objective that introduces latent feedback late in pretraining and mixes a small fraction of deeper feedback passes for stability. We train 1B-parameter full-bandwidth transformers up to 400B tokens and find that latent feedback improves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance. With negligible per-token decoding overhead, full-bandwidth transformers match or approach standard transformers trained with roughly $1.5\times$ more tokens, and manage to produce shorter reasoning traces at equal or better accuracy.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.