Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
AuthorsKevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
Resources
This study finds that chain-of-thought text can hint at which steps matter, but remains far from a reliable explanation of how models actually reason.
Key results
Qwen3 responses across six mathematical reasoning benchmarks.
Rollouts used to estimate value and advantage for each reasoning step.
Absolute advantage required for a step to be labeled consequential.
Responses with at least one consequential step under a cue, versus 58% without the cue.
Upper end of the 0.28–0.30 in-distribution range for fine-tuned critics.
Upper end of the 0.065–0.10 range, indicating weak recovery of important steps in correct traces.
What the paper found
This paper argues that legible chain-of-thought is not automatically interpretable: a reasoning step’s functional importance should be measured as its reinforcement-learning advantage, the change in expected probability of reaching the model’s final answer after that step. Using Monte Carlo rollouts, the Pruned Exact Linear Time changepoint algorithm, and an effect threshold of 0.1, the study labels consequential steps across 1800 Qwen3 responses spanning AIME 24, AIME 25, AIME 26, AMC 23, MATH500, and GSM8K, with 50 rollouts per prefix. Thinking and model scale mainly increase the share of responses already likely to be correct before reasoning begins: the “high throughout” pattern rises from 24% to 61% with thinking and from 25% to 39% across model sizes. In a Scruples cue experiment, consequential responses fall from 58% without a cue to 15% with one, suggesting that apparently rich reasoning can become behaviorally redundant. Text-only judges and fine-tuned critics recover importance only partially: critics reach PR-AUC 0.28–0.30 on incorrect responses in-distribution but just 0.065–0.10 on correct responses, far below the noise ceiling. The findings challenge process reward models and LLM judges that treat semantic fluency as evidence of causal contribution. The experiments center on Qwen3, with Cohere-linked development and coding assistance from Claude and GPT-5-Codex, rather than systems such as ChatGPT or DeepSeek.
Original abstract
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.