Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
AuthorsYing Fan, Anej Svete, Kangwook Lee
LOTUS uses recurrent Transformer computation to make hidden-state reasoning faster while retaining much of the power and interpretability of explicit chain-of-thought.
Key results
GSM8K-Aug training examples
Llama-3.2-3B-Instruct main scaling model
LOTUS accuracy on the GSM8K test set with Llama-3.2-3B-Instruct
LOTUS speedup over explicit CoT on compact mathematical traces
LOTUS speedup over explicit CoT on natural-language CoT
Gold intermediate-token recovery from post-loop latents
What the paper found
Researchers from Microsoft Research, ETH Zürich, KRAFTON, and Ludo Robotics introduce LOTUS, a latent chain-of-thought method designed to retain explicit reasoning accuracy without decoding every intermediate token. Motivated by the test-time-compute trend associated with systems such as OpenAI’s reasoning models and DeepSeek, LOTUS combines a looped Transformer with a padded latent workspace: six latent blocks of 25 positions are refined for six shared-weight iterations, while parallel cross-entropy directly aligns each latent position with its gold chain-of-thought token. Experiments train on GSM8K-Aug’s 385k examples and evaluate GPT-2, Llama-3.2-1B-Instruct, and Llama-3.2-3B-Instruct. On Llama-3.2-3B-Instruct, LOTUS reaches 70.0% GSM8K accuracy, versus 71.5% for explicit CoT, and achieves a 63.9% out-of-domain average across GSM-Hard, MultiArith, and SVAMP, exceeding explicit CoT’s 62.1%. Its thought phase is 2.5-fold faster than explicit CoT for compact mathematical traces; on natural-language reasoning, it attains 68.13% versus 68.41% for explicit CoT while delivering a 6.9-fold speedup. Ablations show that both recurrent depth and direct gold-token supervision are necessary: increasing loop depth to six raises accuracy to 70.0%, while projecting final latents through the language-model head recovers 70.9% of gold CoT tokens at top-1 and surfaces unseen but valid alternatives, supporting the claim that LOTUS’s latent space is interpretable and CoT-aligned.
Original abstract
Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent CoT offers an alternative: it performs multi-step reasoning in the model's hidden states, replacing decoded tokens with continuous representations for greater efficiency. However, existing latent CoT methods underperform explicit CoT beyond 1B parameters, and the gap widens with scale. Looped, or recurrent-depth, Transformers, which reuse their weights to increase computation depth without adding parameters, are a natural fit for latent reasoning. We therefore ask whether looped Transformers can bridge this gap. We answer affirmatively with a simple recipe: a looped padded Transformer that processes K latent blocks in parallel for R iterations, with a cross-entropy loss on each latent position's gold CoT-step token, similar to explicit CoT supervision. We instantiate it as LOTUS (Looped Transformers with parallel supervision on latents). LOTUS is, to our knowledge, the first latent-CoT method to bridge the gap to explicit CoT at the 3B scale, while cutting thought-phase latency by 2.5x-6.9x from compact math expressions to natural language. Projecting LOTUS's post-loop latents through the base LM head recovers the gold reasoning steps and even surfaces alternative valid intermediate steps, evidence that its latent space is interpretable and CoT-aligned. Ablations confirm that both the looped backbone and the parallel supervision on gold CoT tokens are essential. Code is available at https://github.com/yingfan-bot/lotus.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.