From Base Rollouts to RL Reasoning: A Budgeted Search Perspective
AuthorsWenhe Sun, Cunxiang Wang, Zijun Yao, Yixin Cao
Resources
The study argues that much of RL’s apparent reasoning improvement may come from making models sample promising existing reasoning paths more efficiently.
Key results
Mean percentage-point error on Qwen2.5-7B across Math500, AIME, GPQA, and IFEval.
Mean percentage-point error, reported as 3.07 ± 0.39 pp.
Mean error on AMC23, Minerva, OlympiadBench, and TinyMMLU.
Overall transfer error without target-model RL calibration.
19 of 23 model-benchmark cells recover RL self-consistency behavior.
What the paper found
This paper investigates whether reinforcement learning with verifiable rewards creates new reasoning capabilities or mainly makes existing base-model reasoning easier to sample. Motivated by systems such as OpenAI o1 and DeepSeek-R1, it introduces SearchLens, a Unified Decoding Framework that represents temperature sampling, beam search, tree search, and resampling as policies operating under explicit inference budgets, while evaluating outputs separately with pass@k, self-consistency, best-of-N, and first-finish success. Using paired Base/RL checkpoints from SimpleRL-Zoo, including Qwen2.5-7B, the proposed BOPTR-P1 rule traces an RL model’s default pass@k curve through structured base-model operating points with NBase ≈ αNRL^β. The exponent is benchmark-dependent: β equals 0.60 for Math500, 1.00 for GPQA and IFEval, and 0.00 for AIME. On the Qwen2.5-7B anchor, BOPTR-P1 reaches 3.41 pp mean transfer error, with a three-seed replication at 3.07 ± 0.39 pp. Across ten models and four model families, new-checkpoint errors range from 3.28 to 4.87 pp, while four held-out benchmarks average 5.03 pp. A base-only model-offset predictor achieves 4.19 pp without target-model RL calibration, and 83 percent of self-consistency cells are recoverable. The results support a qualified internalized-search interpretation: much of the tested RL gain reflects improved sampling efficiency toward trajectories already accessible to the base model, but transfer degrades across families, tasks, and protocols, so BOPTR is a behavioral diagnostic rather than evidence of parameter-level equivalence or a universal predictor of post-RL performance.
Original abstract
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but how these gains relate to inference-time decoding and search remains unclear. Does RL create reasoning the base model lacks, or shift the rollout distribution toward trajectories it can already reach but rarely samples? We study this behaviorally with a Unified Decoding Framework (UDF), which expresses token-level sampling, beam-like search, tree search, and sequence-level resampling as executable policies over a shared budgeted operating space, scored post hoc with pass@$k$, self-consistency, best-of-$N$, and first-finish success. Using paired Base/RL checkpoints from SimpleRL-Zoo, we ask whether an RL default-policy curve can be approximated by a structured path of Base operating points. On Math500, AIME, GPQA, and IFEval, the pass@$k$ recovery path follows a Budgeted Operating-Point Transition Rule (BOPTR), $N_{\mathrm{Base}} \approx αN_{\mathrm{RL}}^β$, with benchmark-conditioned exponents. On Qwen2.5-7B, BOPTR gives the lowest transfer error among the non-oracle rules we test, 3.41 pp (95% CI [2.32, 5.53]); a three-seed replication gives 3.07 $\pm$ 0.39 pp. The rule extends to ten models across four families (3.28 to 4.87 pp on checkpoints added after fitting), to four benchmarks it was never fitted on (5.03 pp vs. 4.44 pp in fit), and holds without an RL checkpoint for the target model (4.19 pp) or without RL supervision of any kind (5.08 pp). These results support a qualified internalized-search reading: under the recipe we test, much of the measured RL gain corresponds to a change in sampling efficiency toward operating points the base model can already reach under search. We treat the scaling patterns as descriptive of this recipe and cohort, report where they break down, and use UDF and BOPTR as behavioral diagnostics rather than evidence of parameter-level equivalence.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.