Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck
AuthorsDavide Romano, Kanak Raj, Jerrod Parker, Daniele Giofrè
Resources
Test-time scaling can generate better answers, but for open-ended tasks the real challenge is reliably choosing or synthesizing the best one.
Key results
Average Spearman correlation between reward-model scores and true quality on open-ended generation, making Best-of-N selection nearly random.
Beam Search and Particle Filtering retain only 40–60% of the diversity achieved by independent sampling.
Overall XHigh score for Fusion with Qwen3.5-35B-A3B across the five open-ended benchmarks.
Overall single-sample inference score used as the baseline for Qwen3.5-35B-A3B.
Approximate fraction of available exploration headroom recovered by Fusion, the strongest tested method.
What the paper found
This study presents a compute-normalized comparison of five test-time scaling families—Best-of-N, Beam Search, Particle Filtering, Sequential Refinement, and Fusion—across HealthBench, PRBench, LEXam, WildBench, and WritingBench, covering medicine, professional reasoning, law, chat, and creative writing. Using Qwen3.5 and OLMo3 generators, with reward models including Skywork-Reward-V2 and Llama-3.1-70B, the researchers separate exploration, which generates candidate answers, from exploitation, which selects or synthesizes the final response. Exploration keeps working: the bias-corrected oracle quality of larger candidate pools rises across benchmarks, but exploitation fails because reward-model rankings correlate with judged quality at only 0.12, making Best-of-N selection nearly random. Beam Search and Particle Filtering worsen this through diversity collapse, retaining just 40–60% of the diversity of independent sampling. Sequential Refinement produces a genuine gain on only one of five benchmarks, while apparent improvements on WritingBench are strongly confounded by verbosity bias; the benchmark’s native judge includes Claude 3.7 Sonnet among its evaluators, while HealthBench originates from OpenAI. Fusion is the only consistently improving strategy for Qwen3.5-35B-A3B, reaching an overall score of 0.610 at the highest compute level versus 0.574 for single-sample inference, yet it captures only 40% of available exploration headroom. Results from OLMo3 show that generating strong candidates does not guarantee the ability to exploit them, establishing candidate selection and synthesis—not candidate generation—as the central bottleneck in open-ended test-time scaling.
Original abstract
Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation. The answer depends on which side of that decomposition you examine. Scaling exploration works: the best candidate in the pool improves steadily with compute across all settings. What breaks is exploitation - the step that converts a rich candidate pool into a final output. With state-of-the-art generators, reward models correlate at only $ρ_v \approx 0.12$ with true quality, rendering selection near-random regardless of budget. Tree search amplifies this failure through diversity collapse. Refinement helps on one of five benchmarks; its apparent gains elsewhere are confounded. Only synthesis across candidates (Fusion) consistently improves over single-sample baselines, yet still recovers only ~40% of available quality. The candidate pool is not the bottleneck - choosing from it is.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.