NTH

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

AuthorsQiancheng Zhou, Ruizhe Li

September 12, 2026 3 min read
Watch on YouTube
The one-line take

RLVR makes models better at one answer by locking them into a narrow path right at the entrance of the reasoning process.

Key results

67%
PPO solution-coverage contraction

Coverage falls from 0.337 to 0.111 at 320 samples while pass@1 increases.

16
PPO entrance likelihood-shift ratio

The pre-operation likelihood shift is 16-fold larger than the downstream execution shift.

0.212
Minimal-entrance completion

Supplying an unselected entrance raises completion from 0.018 to 0.212.

37%
Layer-interpolation coverage gain

Interpolating late layers 20–28 with early-checkpoint weights increases solution coverage without pass@1 loss.

What the paper found

This paper argues that reinforcement learning with verifiable rewards, or RLVR, makes reasoning models more accurate by locking them into a narrow set of opening strategies. Using exhaustively enumerated Countdown solutions, the researchers separate access—the decision to enter a valid solution family—from execution, the arithmetic that follows. In PPO training of Qwen2.5-3B, solution coverage at 320 samples falls from 0.337 to 0.111, a 67% contraction, even as pass@1 rises sharply; the same pattern appears with GRPO on Qwen2.5-3B-Instruct. Teacher-forced likelihood shifts are 16-fold larger before the first arithmetic operation than during downstream reasoning under PPO, and 11-fold larger under GRPO, indicating that RLVR suppresses alternative entrances rather than erasing arithmetic capability. Supplying only an otherwise unselected operand-and-operator prefix raises low-access family completion from 0.018 to 0.212, showing that forgotten paths remain executable. Surface prompts and temperature changes recover little diversity, but interpolating late layers 20–28 with early-checkpoint weights increases solution coverage by 37% without reducing pass@1. The entrance-collapse signature generalizes across six benchmarks, including GSM8K, MATH500, AMC23, and AIME24, on 7B and 14B models. Crucially, staged SFT–DPO–RLVR training, as in OLMo-3, preserves early-step entropy, while DeepSeek-R1-Distill-Qwen-7B retains both high accuracy and broad entrances, suggesting that narrow search is a consequence of on-policy optimization rather than reasoning itself.

Original abstract

Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator, across PPO on Qwen2.5-3B and GRPO on Qwen2.5-3B-Instruct. Across both training setups, solution coverage falls by up to 67%, halving even on problems solved across all checkpoints. We show that this contraction is heavily concentrated at the entrance: per-token likelihood shifts are 11x--16x larger prior to the first arithmetic operation than during downstream reasoning. Supplying only an unselected entrance prefix restores completion rates in low-access families by over an order of magnitude (0.018 -> 0.212 under PPO), demonstrating that alternative solutions remain executable but are no longer initiated. Guided by this localization, we find that while surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1. Finally, we show that early-step entropy collapse recurs across six math benchmarks with 7B and 14B models, but is not an inevitable byproduct of reasoning optimization: an SFT baseline preserves more than double the coverage, and staged SFT--DPO--RLVR pipelines retain early-step entropy. In summary, reasoning breadth is lost at the door, not inside the room. Code: https://github.com/ershiyidian/early-branch-locking.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →