NTH

Verifier-Induced Support Reshaping in On-Policy Optimization

AuthorsShaohang Wei, Zikun Su, Feifan Song, Wen Luo, Wei Li, Guangyue Peng, Houfeng Wang

August 18, 2026 3 min read
Watch on YouTube
The one-line take

RLVR may make a model better at today’s task while quietly making tomorrow’s successful behaviors harder to discover.

Key results

6.5
Qwen3-8B-Base IFEval pass@1 change

Math-RLVR increases pass@1 by 6.5 percentage points.

9.8
Qwen3-8B-Base IFEval best@32 change

Math-RLVR decreases best@32 by 9.8 percentage points.

32
Core rollout budget

Support evaluation uses 32 stochastic rollouts per prompt.

106.7
Qwen2.5-Math-7B AIME opening divergence ratio

First-token JS divergence is 106.7× the interior divergence under IF-RLVR.

1.6%
IF-first mixed support at step 20

Mixed-support share falls from 33.6% to 1.6% after 20 Math-RLVR steps.

0.0879
Converged-teacher OPD MATH-500-128 mean@16

OPD reduces MATH-500-128 mean@16 to 0.0879 from 0.3433.

What the paper found

This paper identifies verifier-induced support reshaping in on-policy reinforcement learning with verifiable rewards, or RLVR: optimizing one objective can make successful behaviors for a later objective too rare to sample within a fixed rollout budget. Experiments with Qwen3-8B-Base and Qwen2.5-Math-7B compare Math-RLVR on the 7.5k MATH split with IF-RLVR on IFTrain, using 32 stochastic rollouts per prompt and evaluating AIME, IFEval, IFBench, MathIF, and ReasonIF. On IFEval with Qwen3-8B-Base, Math-RLVR increases pass@1 by 6.5 percentage points but decreases best@32 by 9.8 percentage points, showing higher average success alongside narrower prompt coverage. Conversely, IF-RLVR shifts mathematical responses from deliberative-reasoning initiation to direct-answer initiation and reduces math searchability. The largest distributional change occurs at the first generated token: on AIME with Qwen2.5-Math-7B, first-token Jensen–Shannon divergence is 106.7× the interior divergence. For IF-first sequential training, the mixed-support share falls from 33.6% to 1.6% by step 20, leaving little reward variation for later Math-RLVR. Routing priors and on-policy distillation provide only partial protection; distillation from a converged IF teacher reduces MATH-500-128 mean@16 from 0.3433 to 0.0879. An independent DeepSeek-V4-Pro audit also finds verifier-passed shortcut responses reaching 73.5%, emphasizing that endpoint scores can conceal degraded searchability and response quality. The central conclusion is that RLVR mainly reranks openings already available in the base policy, so preserving future trainability requires tracking effective support, not just current benchmark accuracy.

Original abstract

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. We call this verifier-induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Across two model families, we study this effect through repeated verifier-scored sampling and bidirectional training on mathematical reasoning and constrained instruction following, including sequential training with the opposite verifier. Math-RLVR raises average instruction-following success but reduces the number of prompts with any successful response under repeated sampling. On IFEval with Qwen3-8B-Base, pass@1 rises by 6.5 percentage points while best@32 falls by 9.8 percentage points, and the same divergence appears across both models and IF benchmarks. Conversely, IF-RLVR shifts math responses from step-by-step openings toward direct answers, lowers best@k across sampling budgets, and reduces reward variation for later Math-RLVR. Token-distribution analyses and controlled opening interventions show that these changes concentrate in the first few response tokens. RLVR mainly reranks openings already available in the base policy, and the selected opening causally affects math searchability. The tested reference-policy constraints, routing priors, and on-policy distillation preserve cross-task support only partially; MathIF and ReasonIF show that marginal gains translate only partly into responses that are both correct and constraint-following. Therefore, endpoint improvements do not guarantee future trainability or joint capability under on-policy optimization. Code is available at https://github.com/sylvain-wei/verifier-induced-support-reshaping

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →