Principled Thoughts for Latent Recursive LLM Systems
AuthorsFahd Seddik, Fatemeh Fard
AffiliationsFARD Lab, University of British Columbia, Okanagan, Canada
Resources
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.
Key results
REST was tested across 7 mathematical, scientific, medical, and code-generation benchmarks.
Average accuracy improvement in percentage points over CE-only training.
Average accuracy improvement in percentage points over CE-only training.
Maximum accuracy improvement in percentage points.
REST’s final-answer convergence rate, compared with 73% for CE-only.
Average increase in decoded tokens under REST.
What the paper found
The paper introduces REST, or REpresentation-Supervised Thoughts, for latent recursive LLM systems that reason through hidden states rather than decoded chain-of-thought text. It identifies four failures of training only with final-answer cross-entropy: latent thoughts can lose causal information, preserve irrelevant input, collapse across distinct questions, and encode one sampled output instead of uncertainty. REST adds differentiable losses for causality, minimality, separability, and stability to the existing objective, training only the outer communication link while leaving the base models and inference architecture unchanged. Experiments use Qwen, Llama, and Gemma models in single-agent self-recursion and planner–refiner–solver multi-agent systems, evaluated across 7 benchmarks including MATH500, GPQA-Diamond, MedQA, AIME2025, AIME2026, LiveCodeBench-v6, and MBPP+. Under matched data, compute, and latent budgets, REST improves average accuracy by 3.3 percentage points for single-agent systems and 3.5 percentage points for multi-agent systems, with best gains reaching 6.5 and 7.5 percentage points. The method also produces more separable and output-focused thoughts, recovers 65% of oracle-text accuracy compared with 34% for CE-only transfer, and raises the boxed-answer rate from 73% to 95%. REST decodes 15.4% more tokens on average, but the added computation is associated with more frequent convergence and, in some cases, less repetitive reasoning.
Original abstract
Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning
Shuvendu K. Lahiri
NFV uses AI agents to translate ordinary code into machine-checkable formal proofs, making software verification more accessible while revealing the limits of end-to-end soundness.