Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning
AuthorsSangmook Lee, Minbeom Kim, Jeonghye Kim, Dohyung Kim, Sojeong Rhee, Kyomin Jung
Resources
This paper asks whether LLMs are truly solving math problems in different ways, or just wording the same strategy differently, and finds that many diversity metrics miss the real story.
Key results
Evaluation problems retained with at least three distinct correct approaches.
GPT-5.2 judge agreement with the human approach-diversity reference.
Median overlapping-unigram ratio among correct solutions.
Share of DIVER’s textual-diversity improvement attributable to same-approach variation.
Approach-coverage decline after optimizing the in-loop Qwen judge reward.
Coverage decline measured by the judge used during training.
What the paper found
Researchers at Seoul National University and KAIST argue that most diversity metrics for LLM mathematical reasoning measure phrasing rather than strategy. They define approach-level diversity as variation in mathematical tools, problem structure, or representational viewpoint, and evaluate it with a human-calibrated GPT-5.2 judge. On 80 solution pairs from the MATH dataset, human annotators agreed on 80% of cases, while the judge reached 85.0% agreement. Scaling the framework to 469 MATH problems showed that lexical, embedding, symbolic, and reasoning-path metrics can recognize coarse differences but fail to detect incremental strategic diversity once both candidate sets already contain multiple approaches. Shared scaffolding was substantial, with median token overlap of 61.2%, and paraphrasing could make identical strategies appear more diverse than genuinely different solutions. In diversity-aware RLVR experiments with Qwen2.5 models, DQO and DIVER preserved their optimized surface metrics but reduced approach coverage; notably, 80.6% of DIVER’s textual-diversity gain came from variation within the same approach. Approach-controlled test-time scaling showed that covering more strategies improves self-consistency, best-of-N, and pass@k, across Qwen2.5-3B-Instruct, Meta’s Llama3.2-3B-Instruct, and OpenAI’s GPT-4o-mini. Directly rewarding judge-assessed diversity also failed: after training, coverage measured by an external GPT judge fell 38%, versus an 18% decline under the in-loop Qwen judge, indicating reward hacking rather than broader reasoning. The paper concludes that robust, training-compatible measurement of genuine strategic diversity remains open.
Original abstract
Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than differences in how a problem is solved. We address this gap by introducing approach-level diversity: variation in strategies across correct solutions to the same problem. Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while approach-level diversity declines. Investigating when approach-level diversity helps and whether it can be directly induced, we find that approach-diverse candidate sets improve test-time scaling. However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem. Together, our work introduces the notion of approach-level diversity and uncovers a systematic divergence between surface- and approach-level signals, marking a step toward LLMs that reason in genuinely diverse, human-like ways.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.