Trading Depth for Time in Recurrent Transformers
AuthorsZeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen
AffiliationsMicrosoft · University of Wisconsin–Madison
Resources
This paper shows that letting a Transformer spend extra computation on latent thought steps can approach the quality of doubling its depth while using substantially fewer parameters.
Key results
BPB for the 16-layer LRT with one thought token.
BPB for the 20-layer LRT with one thought token.
Fraction of the BPB improvement from doubling depth recovered by one thought token.
Fraction of the BPB improvement from doubling depth recovered by one thought token.
Fewer total parameters than the corresponding double-depth LRT.
BPB degradation after removing explicit cross-token recurrence from the 16-layer one-thought model.
What the paper found
Trading Depth for Time in Recurrent Transformers asks whether recurrent models benefit more from adding independently parameterized layers or from inserting extra computation steps in time. The method extends Latent Recurrent Transformers, or LRTs, with continuous thought tokens placed between vocabulary tokens; each thought token reuses the same backbone, hidden-state feedback, KV feedback, attention, and mixture-of-experts routing before the next token is predicted. On Microsoft’s NanoChat MoE backbones, a 16-layer LRT with one thought token reaches 0.786 BPB versus 0.780 for a 32-layer LRT, while the 20-layer version reaches 0.745 versus 0.741 for a 40-layer model. These temporal variants recover 66.7% and 81.0% of the BPB improvement from doubling depth while using 48.0% and 48.5% fewer total parameters, with both alternatives executing two backbone passes per vocabulary token. Adding a second thought token improves BPB to 0.779 and 0.739 at the two reference sizes, although triple-depth LRTs remain better at the same block count. The key architectural finding is that feeding the final thought state into the next vocabulary token preserves a longer recurrent dependency chain: removing this cross-token recurrence worsens the 16-layer result by 0.009 BPB, while removing KV feedback worsens it by 0.004. Compared with PonderLM-2 and looped Transformers, LRT’s explicit cross-token feedback delivers stronger parameter-efficient gains, though physical depth still provides superior capacity.
Original abstract
Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same $L$ layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this $L$-layer LRT against a $2L$-layer LRT without thought tokens. Both execute $2L$ Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.