NTH

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

AuthorsMohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary

August 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains how to fairly compare reasoning LLMs that spend more computation at inference time and provides practical standards for evaluating and reproducing them.

Key results

3
Structural test-time-scaling regimes

The taxonomy covers single-trajectory, leaf-level, and prefix-level inference.

1,948,821
Released reasoning traces

Full traces span broad knowledge, symbolic reasoning, and competition mathematics.

91.67%
Qwen3 Pass@80

Qwen3-30B-A3B-Thinking-2507’s candidate-bank discovery rate on repeated mathematics.

43.33%
Qwen3 all-correct@80

The probability that all 80 sampled responses are correct for Qwen3-30B-A3B-Thinking-2507.

86.56%
Qwen3.6 reference-free selection

Pointwise Best-of-N accuracy versus 94.62% Pass@80 on competition mathematics.

What the paper found

This paper reframes test-time scaling in reasoning LLMs as budgeted inference over an autoregressive model’s implicit prefix tree, distinguishing 3 regimes: single-trajectory sequential deliberation, leaf-level sampling with terminal reduction, and prefix-level search over unfinished states. It argues that the evaluated object is the complete inference system—checkpoint, prompt, decoder, controller, verifier or judge, stopping rule, budget, and uncertainty procedure—not model weights alone. Its evaluation framework separates end-to-end utility from candidate-bank diagnostics through a discovery–stability profile that unifies Pass@k, all-correct rates, and majority-style metrics, while requiring compute accounting across generation, evaluation, control, and decision costs. Reproducibility is divided into exact replay and distributional reproducibility, with released prompts, completions, metadata, verifier records, random streams, and protocol-matched bootstrap uncertainty. The accompanying corpus contains 1,948,821 full reasoning traces. Experiments across MMLU-Pro, BIG-Bench Hard, AIME, HMMT, BrUMO, MathArena, and SuperGPQA show why reducers matter: for Qwen3-30B-A3B-Thinking-2507, Pass@80 reaches 91.67% while all-correct@80 is 43.33%, and mean-log-probability Best-of-N falls from 75.56% to 65.83% as the bank grows. On competition mathematics, Qwen3.6 reaches 94.62% Pass@80 but only 86.56% with a reference-free pointwise verifier. The findings apply directly to systems such as DeepSeek-R1, Qwen3, Microsoft Phi-4-reasoning, and OpenAI gpt-oss-20b: more inference compute can expose correct candidates without ensuring that the deployed selector chooses them.

Original abstract

Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate them through voting or verification, or search over unfinished partial states. These algorithms differ in their statistical structure, compute accounting, and failure modes. Treating these procedures as interchangeable under a single scalar "budget," or reporting accuracy without the inference protocol that produced it, makes results difficult to compare across studies. We develop a systematic account of test-time scaling along three axes. First, we formalize test-time scaling as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish three structural regimes: single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the evaluated object as the entire inference system and develop evaluation principles that separate end-to-end system performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and prescribe protocol-matched reporting of compute and uncertainty. Third, we specify reproducibility requirements for inference protocols, distinguishing exact replay from distributional reproducibility and identifying the artifacts needed to support each. We also organize the open-weight reasoning ecosystem by model-side and interface mechanisms, apply these principles to broad-knowledge, symbolic-reasoning, and competition-mathematics benchmarks, and assemble over 2 billion full reasoning traces for release with progressively richer verifier and token-level signals.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis