Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
AuthorsJean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig
Resources
The paper finds that thinking models often amplify the appearance of careful reasoning more than the behaviors that truly signal correct answers.
Key results
Reasoning traces analyzed across the study
Language and vision-language models compared
Evaluation benchmarks spanning logical, mathematical, visual, and knowledge tasks
Behavioral Lift associated with confidence calibration in language models
Negative Behavioral Lift magnitude for uncertainty acknowledgment
Recovery rate for thinking models after detected failures
What the paper found
This study tests whether the longer, more deliberative traces produced by thinking models such as OpenAI o1, DeepSeek-R1, and Qwen3 actually contain the behaviors most predictive of correct answers. Using Behavioral Lift—defined as the accuracy difference when a behavior is present versus absent—researchers analyzed 15,282 traces from 15 language and vision-language models across 6 benchmarks, including MATH-500, MMLU-Pro, MathVista, and LogiQA2, with GPT-4o providing taxonomy annotations. Thinking models substantially amplify self-correction, hypothesis testing, and uncertainty acknowledgment, but these are not the strongest correctness signals. Confidence calibration, knowledge alignment, and self-awareness rank higher: confidence calibration has 79.6% lift for language models, while uncertainty acknowledgment has negative 13.9% lift, meaning explicit hesitation is more associated with failure than success. The main advantage of thinking models comes from recovery after intermediate errors: on MATH-500, their recovery rate reaches 40.8%, compared with 17.8% for instruct models. The pattern reverses on LogiQA2, where instruct models score 58.4% versus 54.1% for thinking models because rapid logical-form recognition is more useful than extended search. The findings argue for process objectives that reward calibrated confidence, evidence grounding, domain alignment, and effective recovery—not merely longer chains of thought or visible hesitation.
Original abstract
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.