Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
AuthorsMohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary
Resources
This paper explains how to fairly compare reasoning LLMs that spend more computation at inference time and provides practical standards for evaluating and reproducing them.
Key results
The taxonomy covers single-trajectory, leaf-level, and prefix-level inference.
Full traces span broad knowledge, symbolic reasoning, and competition mathematics.
Qwen3-30B-A3B-Thinking-2507’s candidate-bank discovery rate on repeated mathematics.
The probability that all 80 sampled responses are correct for Qwen3-30B-A3B-Thinking-2507.
Pointwise Best-of-N accuracy versus 94.62% Pass@80 on competition mathematics.
What the paper found
This paper reframes test-time scaling in reasoning LLMs as budgeted inference over an autoregressive model’s implicit prefix tree, distinguishing 3 regimes: single-trajectory sequential deliberation, leaf-level sampling with terminal reduction, and prefix-level search over unfinished states. It argues that the evaluated object is the complete inference system—checkpoint, prompt, decoder, controller, verifier or judge, stopping rule, budget, and uncertainty procedure—not model weights alone. Its evaluation framework separates end-to-end utility from candidate-bank diagnostics through a discovery–stability profile that unifies Pass@k, all-correct rates, and majority-style metrics, while requiring compute accounting across generation, evaluation, control, and decision costs. Reproducibility is divided into exact replay and distributional reproducibility, with released prompts, completions, metadata, verifier records, random streams, and protocol-matched bootstrap uncertainty. The accompanying corpus contains 1,948,821 full reasoning traces. Experiments across MMLU-Pro, BIG-Bench Hard, AIME, HMMT, BrUMO, MathArena, and SuperGPQA show why reducers matter: for Qwen3-30B-A3B-Thinking-2507, Pass@80 reaches 91.67% while all-correct@80 is 43.33%, and mean-log-probability Best-of-N falls from 75.56% to 65.83% as the bank grows. On competition mathematics, Qwen3.6 reaches 94.62% Pass@80 but only 86.56% with a reference-free pointwise verifier. The findings apply directly to systems such as DeepSeek-R1, Qwen3, Microsoft Phi-4-reasoning, and OpenAI gpt-oss-20b: more inference compute can expose correct candidates without ensuring that the deployed selector chooses them.
Original abstract
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate them through voting or verification, or search over unfinished partial states. These algorithms differ in their statistical structure, compute accounting, and failure modes. Treating these procedures as interchangeable under a single scalar "budget," or reporting accuracy without the inference protocol that produced it, makes results difficult to compare across studies. We develop a systematic account of test-time scaling along three axes. First, we formalize test-time scaling as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish three structural regimes: single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the evaluated object as the entire inference system and develop evaluation principles that separate end-to-end system performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and prescribe protocol-matched reporting of compute and uncertainty. Third, we specify reproducibility requirements for inference protocols, distinguishing exact replay from distributional reproducibility and identifying the artifacts needed to support each. We also organize the open-weight reasoning ecosystem by model-side and interface mechanisms, apply these principles to broad-knowledge, symbolic-reasoning, and competition-mathematics benchmarks, and assemble over 2 billion full reasoning traces for release with progressively richer verifier and token-level signals.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.