NTH

Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

AuthorsDavide Romano, Kanak Raj, Jerrod Parker, Daniele Giofrè

August 23, 2026 2 min read
Watch on YouTube
The one-line take

Test-time scaling can generate better answers, but for open-ended tasks the real challenge is reliably choosing or synthesizing the best one.

Key results

0.12
Reward-model quality correlation

Average Spearman correlation between reward-model scores and true quality on open-ended generation, making Best-of-N selection nearly random.

40–60%
Tree-search diversity retention

Beam Search and Particle Filtering retain only 40–60% of the diversity achieved by independent sampling.

0.610
Fusion overall score

Overall XHigh score for Fusion with Qwen3.5-35B-A3B across the five open-ended benchmarks.

0.574
Single-sample baseline

Overall single-sample inference score used as the baseline for Qwen3.5-35B-A3B.

40%
Fusion headroom capture

Approximate fraction of available exploration headroom recovered by Fusion, the strongest tested method.

What the paper found

This study presents a compute-normalized comparison of five test-time scaling families—Best-of-N, Beam Search, Particle Filtering, Sequential Refinement, and Fusion—across HealthBench, PRBench, LEXam, WildBench, and WritingBench, covering medicine, professional reasoning, law, chat, and creative writing. Using Qwen3.5 and OLMo3 generators, with reward models including Skywork-Reward-V2 and Llama-3.1-70B, the researchers separate exploration, which generates candidate answers, from exploitation, which selects or synthesizes the final response. Exploration keeps working: the bias-corrected oracle quality of larger candidate pools rises across benchmarks, but exploitation fails because reward-model rankings correlate with judged quality at only 0.12, making Best-of-N selection nearly random. Beam Search and Particle Filtering worsen this through diversity collapse, retaining just 40–60% of the diversity of independent sampling. Sequential Refinement produces a genuine gain on only one of five benchmarks, while apparent improvements on WritingBench are strongly confounded by verbosity bias; the benchmark’s native judge includes Claude 3.7 Sonnet among its evaluators, while HealthBench originates from OpenAI. Fusion is the only consistently improving strategy for Qwen3.5-35B-A3B, reaching an overall score of 0.610 at the highest compute level versus 0.574 for single-sample inference, yet it captures only 40% of available exploration headroom. Results from OLMo3 show that generating strong candidates does not guarantee the ability to exploit them, establishing candidate selection and synthesis—not candidate generation—as the central bottleneck in open-ended test-time scaling.

Original abstract

Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation. The answer depends on which side of that decomposition you examine. Scaling exploration works: the best candidate in the pool improves steadily with compute across all settings. What breaks is exploitation - the step that converts a rich candidate pool into a final output. With state-of-the-art generators, reward models correlate at only $ρ_v \approx 0.12$ with true quality, rendering selection near-random regardless of budget. Tree search amplifies this failure through diversity collapse. Refinement helps on one of five benchmarks; its apparent gains elsewhere are confounded. Only synthesis across candidates (Fusion) consistently improves over single-sample baselines, yet still recovers only ~40% of available quality. The candidate pool is not the bottleneck - choosing from it is.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis