An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
AuthorsMingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan
Resources
This paper shows that large reasoning models can often solve problems yet still be surprisingly bad at judging whether a step-by-step explanation is actually valid.
Key results
perturbed math problem-solution pairs with valid answers but invalid reasoning
GPT 5.4 evaluation accuracy on VAIR
Gemini 3.1 Pro evaluation accuracy on VAIR
human solving accuracy on the GSM8K-derived study subset
human grading accuracy on VAIR
Qwen2.5-Math-PRM-7B evaluation accuracy on VAIR
What the paper found
This paper exposes a production-evaluation gap in large reasoning models using the Valid-Answer-Invalid-Reasoning, or VAIR, benchmark, a 1,001-item math dataset built by perturbing GSM8K, MATH, and ProcessBench solutions so the final answer stays correct while the reasoning becomes flawed through Missing Premises, Missing Reasoning, Shuffled Reasoning, or Circular Reasoning. Across six frontier models—Claude Sonnet 4.6, Claude Opus 4.7, DeepSeek R1, GPT 5, GPT 5.4, and Gemini 3.1 Pro—solution production remains near-perfect at 94.7% or above, and evaluation is high on control sets, but collapses on VAIR to as low as 47.9% for GPT 5.4 and 52.5% for GPT 5, with even Gemini 3.1 Pro falling to 78.6%. Humans show a far smaller gap: 80.8% on solving versus 74.5% on VAIR, a maximum difference of 6.3%, and their evaluation is only modestly harder than production. The authors attribute the LRM failure to answer confirmation bias: chain-of-thought often reveals independent re-solving followed by blind endorsement or forced rationalization, linear probes show that valid final answers override internal representations of invalid reasoning, and causal patching of answer-token activations flips verdicts and shifts models toward step tracing. They also find that a process reward model, Qwen2.5-Math-PRM-7B, reproduces the same weakness, scoring 67.8% on VAIR despite 93.8% on IAIR and 79.3% on VAVR, suggesting that outcome-focused supervision leaves reasoning evaluation brittle even when models can generate strong proofs and solutions.
Original abstract
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning to solve complex problems. How then do LRMs perform at evaluating reasons? We investigate this with the Valid-Answer-Invalid-Reasoning (VAIR) dataset: math problems and solutions with trivial reasoning flaws but valid answers, designed to isolate reasoning evaluation from the confound of reasoning production. Unlike humans, who we find are only 6% worse at grading than solving such problems, we find a substantial production-evaluation gap in LRMs: frontier models score as low as 48% when evaluating VAIR solutions, despite near-perfect solution production. Why this enigma? Through chain-of-thought (CoT) analysis, we find evidence of an answer confirmation bias: LRMs often produce then check for the correct answer instead of carefully verifying each step, fabricating rationalizations even when noticing anomalous reasoning. Linear probes corroborate this, showing that while LRM activations encode some representation of valid reasoning, they fail to robustly represent VAIR solutions as invalid. Causal patching of the final answer's representations causes LRM verdicts and activations to flip, demonstrating that answer validity is responsible for models' confirmation biases. These findings indicate an outstanding limitation in dominant approaches to reasoning training, which incentivize LRMs to produce and confirm reasoning towards correct answers, but not to robustly evaluate the underlying reasons.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.