NTH

Trading Depth for Time in Recurrent Transformers

AuthorsZeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen

AffiliationsMicrosoft · University of Wisconsin–Madison

September 28, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that letting a Transformer spend extra computation on latent thought steps can approach the quality of doubling its depth while using substantially fewer parameters.

Key results

0.786
16-layer one-thought BPB

BPB for the 16-layer LRT with one thought token.

0.745
20-layer one-thought BPB

BPB for the 20-layer LRT with one thought token.

66.7%
16-layer depth-gain recovery

Fraction of the BPB improvement from doubling depth recovered by one thought token.

81.0%
20-layer depth-gain recovery

Fraction of the BPB improvement from doubling depth recovered by one thought token.

48.0%
16-layer parameter reduction

Fewer total parameters than the corresponding double-depth LRT.

0.009
Cross-token recurrence ablation

BPB degradation after removing explicit cross-token recurrence from the 16-layer one-thought model.

What the paper found

Trading Depth for Time in Recurrent Transformers asks whether recurrent models benefit more from adding independently parameterized layers or from inserting extra computation steps in time. The method extends Latent Recurrent Transformers, or LRTs, with continuous thought tokens placed between vocabulary tokens; each thought token reuses the same backbone, hidden-state feedback, KV feedback, attention, and mixture-of-experts routing before the next token is predicted. On Microsoft’s NanoChat MoE backbones, a 16-layer LRT with one thought token reaches 0.786 BPB versus 0.780 for a 32-layer LRT, while the 20-layer version reaches 0.745 versus 0.741 for a 40-layer model. These temporal variants recover 66.7% and 81.0% of the BPB improvement from doubling depth while using 48.0% and 48.5% fewer total parameters, with both alternatives executing two backbone passes per vocabulary token. Adding a second thought token improves BPB to 0.779 and 0.739 at the two reference sizes, although triple-depth LRTs remain better at the same block count. The key architectural finding is that feeding the final thought state into the next vocabulary token preserves a longer recurrent dependency chain: removing this cross-token recurrence worsens the 16-layer result by 0.009 BPB, while removing KV feedback worsens it by 0.004. Compared with PonderLM-2 and looped Transformers, LRT’s explicit cross-token feedback delivers stronger parameter-efficient gains, though physical depth still provides superior capacity.

Original abstract

Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same $L$ layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this $L$-layer LRT against a $2L$-layer LRT without thought tokens. Both execute $2L$ Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis