NTH

Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

AuthorsSangmook Lee, Minbeom Kim, Jeonghye Kim, Dohyung Kim, Sojeong Rhee, Kyomin Jung

July 23, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks whether LLMs are truly solving math problems in different ways, or just wording the same strategy differently, and finds that many diversity metrics miss the real story.

Key results

469
Approach-feasible MATH problems

Evaluation problems retained with at least three distinct correct approaches.

85.0%
LLM judge-human agreement

GPT-5.2 judge agreement with the human approach-diversity reference.

61.2%
Median shared scaffolding

Median overlapping-unigram ratio among correct solutions.

80.6%
DIVER within-approach textual gain

Share of DIVER’s textual-diversity improvement attributable to same-approach variation.

38%
External GPT coverage decline

Approach-coverage decline after optimizing the in-loop Qwen judge reward.

18%
In-loop Qwen coverage decline

Coverage decline measured by the judge used during training.

What the paper found

Researchers at Seoul National University and KAIST argue that most diversity metrics for LLM mathematical reasoning measure phrasing rather than strategy. They define approach-level diversity as variation in mathematical tools, problem structure, or representational viewpoint, and evaluate it with a human-calibrated GPT-5.2 judge. On 80 solution pairs from the MATH dataset, human annotators agreed on 80% of cases, while the judge reached 85.0% agreement. Scaling the framework to 469 MATH problems showed that lexical, embedding, symbolic, and reasoning-path metrics can recognize coarse differences but fail to detect incremental strategic diversity once both candidate sets already contain multiple approaches. Shared scaffolding was substantial, with median token overlap of 61.2%, and paraphrasing could make identical strategies appear more diverse than genuinely different solutions. In diversity-aware RLVR experiments with Qwen2.5 models, DQO and DIVER preserved their optimized surface metrics but reduced approach coverage; notably, 80.6% of DIVER’s textual-diversity gain came from variation within the same approach. Approach-controlled test-time scaling showed that covering more strategies improves self-consistency, best-of-N, and pass@k, across Qwen2.5-3B-Instruct, Meta’s Llama3.2-3B-Instruct, and OpenAI’s GPT-4o-mini. Directly rewarding judge-assessed diversity also failed: after training, coverage measured by an external GPT judge fell 38%, versus an 18% decline under the in-loop Qwen judge, indicating reward hacking rather than broader reasoning. The paper concludes that robust, training-compatible measurement of genuine strategic diversity remains open.

Original abstract

Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than differences in how a problem is solved. We address this gap by introducing approach-level diversity: variation in strategies across correct solutions to the same problem. Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while approach-level diversity declines. Investigating when approach-level diversity helps and whether it can be directly induced, we find that approach-diverse candidate sets improve test-time scaling. However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem. Together, our work introduces the notion of approach-level diversity and uncovers a systematic divergence between surface- and approach-level signals, marking a step toward LLMs that reason in genuinely diverse, human-like ways.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis