NTH

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

AuthorsMingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan

July 3, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that large reasoning models can often solve problems yet still be surprisingly bad at judging whether a step-by-step explanation is actually valid.

Key results

1001
VAIR dataset size

perturbed math problem-solution pairs with valid answers but invalid reasoning

47.9%
frontier LRM VAIR low score

GPT 5.4 evaluation accuracy on VAIR

78.6%
frontier LRM VAIR high score

Gemini 3.1 Pro evaluation accuracy on VAIR

80.8%
human production accuracy

human solving accuracy on the GSM8K-derived study subset

74.5%
human VAIR accuracy

human grading accuracy on VAIR

67.8%
PRM VAIR accuracy

Qwen2.5-Math-PRM-7B evaluation accuracy on VAIR

What the paper found

This paper exposes a production-evaluation gap in large reasoning models using the Valid-Answer-Invalid-Reasoning, or VAIR, benchmark, a 1,001-item math dataset built by perturbing GSM8K, MATH, and ProcessBench solutions so the final answer stays correct while the reasoning becomes flawed through Missing Premises, Missing Reasoning, Shuffled Reasoning, or Circular Reasoning. Across six frontier models—Claude Sonnet 4.6, Claude Opus 4.7, DeepSeek R1, GPT 5, GPT 5.4, and Gemini 3.1 Pro—solution production remains near-perfect at 94.7% or above, and evaluation is high on control sets, but collapses on VAIR to as low as 47.9% for GPT 5.4 and 52.5% for GPT 5, with even Gemini 3.1 Pro falling to 78.6%. Humans show a far smaller gap: 80.8% on solving versus 74.5% on VAIR, a maximum difference of 6.3%, and their evaluation is only modestly harder than production. The authors attribute the LRM failure to answer confirmation bias: chain-of-thought often reveals independent re-solving followed by blind endorsement or forced rationalization, linear probes show that valid final answers override internal representations of invalid reasoning, and causal patching of answer-token activations flips verdicts and shifts models toward step tracing. They also find that a process reward model, Qwen2.5-Math-PRM-7B, reproduces the same weakness, scoring 67.8% on VAIR despite 93.8% on IAIR and 79.3% on VAVR, suggesting that outcome-focused supervision leaves reasoning evaluation brittle even when models can generate strong proofs and solutions.

Original abstract

Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning to solve complex problems. How then do LRMs perform at evaluating reasons? We investigate this with the Valid-Answer-Invalid-Reasoning (VAIR) dataset: math problems and solutions with trivial reasoning flaws but valid answers, designed to isolate reasoning evaluation from the confound of reasoning production. Unlike humans, who we find are only 6% worse at grading than solving such problems, we find a substantial production-evaluation gap in LRMs: frontier models score as low as 48% when evaluating VAIR solutions, despite near-perfect solution production. Why this enigma? Through chain-of-thought (CoT) analysis, we find evidence of an answer confirmation bias: LRMs often produce then check for the correct answer instead of carefully verifying each step, fabricating rationalizations even when noticing anomalous reasoning. Linear probes corroborate this, showing that while LRM activations encode some representation of valid reasoning, they fail to robustly represent VAIR solutions as invalid. Causal patching of the final answer's representations causes LRM verdicts and activations to flip, demonstrating that answer validity is responsible for models' confirmation biases. These findings indicate an outstanding limitation in dominant approaches to reasoning training, which incentivize LRMs to produce and confirm reasoning towards correct answers, but not to robustly evaluate the underlying reasons.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis