NTH

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

AuthorsYing Fan, Anej Svete, Kangwook Lee

July 23, 2026 2 min read
Watch on YouTube
The one-line take

LOTUS uses recurrent Transformer computation to make hidden-state reasoning faster while retaining much of the power and interpretability of explicit chain-of-thought.

Key results

385k
Training dataset

GSM8K-Aug training examples

3B
Backbone scale

Llama-3.2-3B-Instruct main scaling model

70.0%
GSM8K accuracy

LOTUS accuracy on the GSM8K test set with Llama-3.2-3B-Instruct

2.5
Math thought-phase speedup

LOTUS speedup over explicit CoT on compact mathematical traces

6.9
Natural-language thought-phase speedup

LOTUS speedup over explicit CoT on natural-language CoT

70.9%
Gold CoT top-1 readout

Gold intermediate-token recovery from post-loop latents

What the paper found

Researchers from Microsoft Research, ETH Zürich, KRAFTON, and Ludo Robotics introduce LOTUS, a latent chain-of-thought method designed to retain explicit reasoning accuracy without decoding every intermediate token. Motivated by the test-time-compute trend associated with systems such as OpenAI’s reasoning models and DeepSeek, LOTUS combines a looped Transformer with a padded latent workspace: six latent blocks of 25 positions are refined for six shared-weight iterations, while parallel cross-entropy directly aligns each latent position with its gold chain-of-thought token. Experiments train on GSM8K-Aug’s 385k examples and evaluate GPT-2, Llama-3.2-1B-Instruct, and Llama-3.2-3B-Instruct. On Llama-3.2-3B-Instruct, LOTUS reaches 70.0% GSM8K accuracy, versus 71.5% for explicit CoT, and achieves a 63.9% out-of-domain average across GSM-Hard, MultiArith, and SVAMP, exceeding explicit CoT’s 62.1%. Its thought phase is 2.5-fold faster than explicit CoT for compact mathematical traces; on natural-language reasoning, it attains 68.13% versus 68.41% for explicit CoT while delivering a 6.9-fold speedup. Ablations show that both recurrent depth and direct gold-token supervision are necessary: increasing loop depth to six raises accuracy to 70.0%, while projecting final latents through the language-model head recovers 70.9% of gold CoT tokens at top-1 and surfaces unseen but valid alternatives, supporting the claim that LOTUS’s latent space is interpretable and CoT-aligned.

Original abstract

Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent CoT offers an alternative: it performs multi-step reasoning in the model's hidden states, replacing decoded tokens with continuous representations for greater efficiency. However, existing latent CoT methods underperform explicit CoT beyond 1B parameters, and the gap widens with scale. Looped, or recurrent-depth, Transformers, which reuse their weights to increase computation depth without adding parameters, are a natural fit for latent reasoning. We therefore ask whether looped Transformers can bridge this gap. We answer affirmatively with a simple recipe: a looped padded Transformer that processes K latent blocks in parallel for R iterations, with a cross-entropy loss on each latent position's gold CoT-step token, similar to explicit CoT supervision. We instantiate it as LOTUS (Looped Transformers with parallel supervision on latents). LOTUS is, to our knowledge, the first latent-CoT method to bridge the gap to explicit CoT at the 3B scale, while cutting thought-phase latency by 2.5x-6.9x from compact math expressions to natural language. Projecting LOTUS's post-loop latents through the base LM head recovers the gold reasoning steps and even surfaces alternative valid intermediate steps, evidence that its latent space is interpretable and CoT-aligned. Ablations confirm that both the looped backbone and the parallel supervision on gold CoT tokens are essential. Code is available at https://github.com/yingfan-bot/lotus.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis