Self-Trained Verification for Training- and Test-Time Self-Improvement
AuthorsChen Henry Wu, Aditi Raghunathan
Resources
This paper shows that teaching a model to verify its own answers better can improve both test-time reasoning loops and training-time self-improvement, leading to stronger performance on hard reasoning tasks.
Key results
STV verifier on Qwen3-8B raises pass@1 on the hardest SciKnowEval split from 1.5% without verification to 21.0%.
Verifier-in-the-loop training improves the generator’s standalone pass@1 by 30% relative even with no verifier at inference.
What the paper found
Self-Trained Verification (STV) addresses a specific bottleneck in reasoning-model self-improvement: verifiers can judge final answers, but they often cannot diagnose plausible wrong solutions well enough to drive refinement. The paper’s key idea is to exploit a supervision asymmetry: when the same model is conditioned on the reference solution, it can usually identify the error, so STV distills this reference-conditioned “teacher” verifier into an unconditioned student using on-policy distillation with an α-divergence objective at α=0.5, plus an RL term for verdict accuracy. Using Qwen3-8B on hard DAPO math problems and SciKnowEval science tasks, STV roughly doubles final-round pass@1 on hard math, and on the hardest scientific reasoning split it jumps from 1.5% without verification to 21.0%, outperforming much larger Qwen3-32B and Qwen3-235B-A22B baselines. The trained verifier also scales better across 20-plus refinement rounds and shows 3–5× higher precision at matched coverage, reducing reward hacking. The second contribution is verifier-in-the-loop training (ViL): after standard RL from verifiable reward has converged, the generator is trained inside the V-R loop against a frozen STV verifier, yielding a further 33% gain in final-round pass@1 and a 30% relative gain in standalone pass@1 with no verifier at inference. This shows that better verification can improve both test-time refinement and training-time generator capability.
Original abstract
Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verification would unlock both, but the capability we want to train, i.e., catching self-generated errors, lacks training signal. To address this challenge, we propose self-trained verification (STV). Our key observation is that, while a model cannot catch these errors alone, it can when shown the reference solution. We turn this asymmetry into a supervision target and train the verifier to imitate a more informed version of itself. At test time, STV substantially improves V-R loops on hard problems, while alternatives (e.g., SFT, RL on verifier scores, and even meta-verifiers) do not. STV roughly doubles accuracy on hard math and lifts it 14x on scientific reasoning tasks (1.5% to 21%). At training time, we additionally train the generator using RL with STV verifier's feedback inside the V-R loop - a procedure we call verifier-in-the-loop training (ViL). Starting from an RL-converged generator, ViL yields a further 33% gain in pass@1. More notably, the generator's standalone pass@1, with no verifier at test time, climbs 30% relative past where standard RL had converged. Hence, the next frontier in reasoning on hard problems may lie in how we train for and with verification.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.