NTH

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

AuthorsYong Yi Bay, Kathleen A. Yearick

July 9, 2026 2 min read
Watch on YouTube
The one-line take

The paper argues that beyond a modest number of samples, test-time scaling for language models stops helping and can even hurt because the real bottleneck is picking the right answer, not generating more candidates.

Key results

1
effective samples ceiling

As n grows, neff saturates at 1/ρ effective draws; the paper states the ceiling generally as 1/ρ.

104
Brown et al. samples per problem

Released repeated-sampling logs used for GSM8K and MATH.

0.4-0.6
estimated correlation range

Between-problem intraclass correlation on the Brown et al. logs.

0.88
coverage on MATH-500 log

Coverage on Beeching et al.’s Llama-3.2-1B-Instruct log.

0.45
self-consistency on MATH-500 log

Plurality-vote plateau on Beeching et al.’s Llama-3.2-1B-Instruct log.

13
median effective answer count

Median distinct answer modes per problem in the MATH-500 log.

What the paper found

This paper, by Yong Yi Bay and Kathleen A. Yearick, argues that test-time sampling in language models has two different ceilings that are often conflated: a correlation ceiling for estimation and a modal ceiling for selection. Reframing repeated decoding as clustered binary sampling, the authors derive the effective sample size neff = n/[1 + (n − 1)ρ], showing that correlated draws saturate at 1/ρ effective samples and that the marginal value of the n-th sample decays quadratically, so extra sampling quickly becomes redundant. They then separate coverage from selection: pass@n can keep rising toward what the model can ever reach, but plurality vote or self-consistency converges to the most common answer, with benchmark-level accuracy capped at the modal-hit rate πmode. On released logs from Brown et al.’s repeated-sampling runs, they estimate between-problem correlation around 0.4–0.6, meaning 104 samples per problem are worth only about 2 independent observations for benchmark-mean estimation. On Beeching et al.’s five-session Llama-3.2-1B-Instruct MATH-500 log, 256 attempts per problem yielded coverage of 0.88 but self-consistency only 0.45, while the median effective answer count was about 13 distinct modes, explaining why more samples sharpened a confident wrong answer rather than improving selection. The practical takeaway is explicit: estimate ρ from logs, report neff alongside n, spend budget on more problems for evaluation, and reduce answer concentration—not raw sample count—when no verifier is available.

Original abstract

People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sampling it many times (test-time scaling), and the more they draw, the more often a correct answer turns up somewhere, so coverage, the fraction of problems with at least one correct try, climbs and appears to be progress. But a deployed system must return one answer, and choosing it, not knowing which try is right, is selection; selection is capped, and past a point extra samples only make the model surer of a confident mistake, even as every draw adds cost. The gap between climbing coverage and stalled selection, the identifiability gap, is the answer a model can produce but not pick. So the real question is not whether to sample but how far, and the answer is: not far. For picking an answer, the vote has already settled within a few dozen draws, the modal ceiling; for scoring a benchmark, sooner still, the correlation ceiling. Beyond that, extra draws cost compute and add nothing, and can even make the answer worse. This paper turns the cutoff into a single number, the effective number of samples, that any sampling run already reveals. The bottleneck is recognizing a right answer, not generating one.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis