NTH

From Base Rollouts to RL Reasoning: A Budgeted Search Perspective

AuthorsWenhe Sun, Cunxiang Wang, Zijun Yao, Yixin Cao

September 4, 2026 3 min read
Watch on YouTube
The one-line take

The study argues that much of RL’s apparent reasoning improvement may come from making models sample promising existing reasoning paths more efficiently.

Key results

3.41
BOPTR-P1 anchor transfer error

Mean percentage-point error on Qwen2.5-7B across Math500, AIME, GPQA, and IFEval.

3.07
Three-seed replication error

Mean percentage-point error, reported as 3.07 ± 0.39 pp.

5.03
Held-out benchmark transfer error

Mean error on AMC23, Minerva, OlympiadBench, and TinyMMLU.

4.19
Base-only offset prediction error

Overall transfer error without target-model RL calibration.

83%
Self-consistency recoverability

19 of 23 model-benchmark cells recover RL self-consistency behavior.

What the paper found

This paper investigates whether reinforcement learning with verifiable rewards creates new reasoning capabilities or mainly makes existing base-model reasoning easier to sample. Motivated by systems such as OpenAI o1 and DeepSeek-R1, it introduces SearchLens, a Unified Decoding Framework that represents temperature sampling, beam search, tree search, and resampling as policies operating under explicit inference budgets, while evaluating outputs separately with pass@k, self-consistency, best-of-N, and first-finish success. Using paired Base/RL checkpoints from SimpleRL-Zoo, including Qwen2.5-7B, the proposed BOPTR-P1 rule traces an RL model’s default pass@k curve through structured base-model operating points with NBase ≈ αNRL^β. The exponent is benchmark-dependent: β equals 0.60 for Math500, 1.00 for GPQA and IFEval, and 0.00 for AIME. On the Qwen2.5-7B anchor, BOPTR-P1 reaches 3.41 pp mean transfer error, with a three-seed replication at 3.07 ± 0.39 pp. Across ten models and four model families, new-checkpoint errors range from 3.28 to 4.87 pp, while four held-out benchmarks average 5.03 pp. A base-only model-offset predictor achieves 4.19 pp without target-model RL calibration, and 83 percent of self-consistency cells are recoverable. The results support a qualified internalized-search interpretation: much of the tested RL gain reflects improved sampling efficiency toward trajectories already accessible to the base model, but transfer degrades across families, tasks, and protocols, so BOPTR is a behavioral diagnostic rather than evidence of parameter-level equivalence or a universal predictor of post-RL performance.

Original abstract

Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but how these gains relate to inference-time decoding and search remains unclear. Does RL create reasoning the base model lacks, or shift the rollout distribution toward trajectories it can already reach but rarely samples? We study this behaviorally with a Unified Decoding Framework (UDF), which expresses token-level sampling, beam-like search, tree search, and sequence-level resampling as executable policies over a shared budgeted operating space, scored post hoc with pass@$k$, self-consistency, best-of-$N$, and first-finish success. Using paired Base/RL checkpoints from SimpleRL-Zoo, we ask whether an RL default-policy curve can be approximated by a structured path of Base operating points. On Math500, AIME, GPQA, and IFEval, the pass@$k$ recovery path follows a Budgeted Operating-Point Transition Rule (BOPTR), $N_{\mathrm{Base}} \approx αN_{\mathrm{RL}}^β$, with benchmark-conditioned exponents. On Qwen2.5-7B, BOPTR gives the lowest transfer error among the non-oracle rules we test, 3.41 pp (95% CI [2.32, 5.53]); a three-seed replication gives 3.07 $\pm$ 0.39 pp. The rule extends to ten models across four families (3.28 to 4.87 pp on checkpoints added after fitting), to four benchmarks it was never fitted on (5.03 pp vs. 4.44 pp in fit), and holds without an RL checkpoint for the target model (4.19 pp) or without RL supervision of any kind (5.08 pp). These results support a qualified internalized-search reading: under the recipe we test, much of the measured RL gain corresponds to a change in sampling efficiency toward operating points the base model can already reach under search. We treat the scaling patterns as descriptive of this recipe and cohort, report where they break down, and use UDF and BOPTR as behavioral diagnostics rather than evidence of parameter-level equivalence.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →