NTH

Self-Trained Verification for Training- and Test-Time Self-Improvement

AuthorsChen Henry Wu, Aditi Raghunathan

May 30, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that teaching a model to verify its own answers better can improve both test-time reasoning loops and training-time self-improvement, leading to stronger performance on hard reasoning tasks.

Key results

21.0%
Hardest SciKnowEval pass@1

STV verifier on Qwen3-8B raises pass@1 on the hardest SciKnowEval split from 1.5% without verification to 21.0%.

30%
ViL standalone pass@1 gain

Verifier-in-the-loop training improves the generator’s standalone pass@1 by 30% relative even with no verifier at inference.

What the paper found

Self-Trained Verification (STV) addresses a specific bottleneck in reasoning-model self-improvement: verifiers can judge final answers, but they often cannot diagnose plausible wrong solutions well enough to drive refinement. The paper’s key idea is to exploit a supervision asymmetry: when the same model is conditioned on the reference solution, it can usually identify the error, so STV distills this reference-conditioned “teacher” verifier into an unconditioned student using on-policy distillation with an α-divergence objective at α=0.5, plus an RL term for verdict accuracy. Using Qwen3-8B on hard DAPO math problems and SciKnowEval science tasks, STV roughly doubles final-round pass@1 on hard math, and on the hardest scientific reasoning split it jumps from 1.5% without verification to 21.0%, outperforming much larger Qwen3-32B and Qwen3-235B-A22B baselines. The trained verifier also scales better across 20-plus refinement rounds and shows 3–5× higher precision at matched coverage, reducing reward hacking. The second contribution is verifier-in-the-loop training (ViL): after standard RL from verifiable reward has converged, the generator is trained inside the V-R loop against a frozen STV verifier, yielding a further 33% gain in final-round pass@1 and a 30% relative gain in standalone pass@1 with no verifier at inference. This shows that better verification can improve both test-time refinement and training-time generator capability.

Original abstract

Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verification would unlock both, but the capability we want to train, i.e., catching self-generated errors, lacks training signal. To address this challenge, we propose self-trained verification (STV). Our key observation is that, while a model cannot catch these errors alone, it can when shown the reference solution. We turn this asymmetry into a supervision target and train the verifier to imitate a more informed version of itself. At test time, STV substantially improves V-R loops on hard problems, while alternatives (e.g., SFT, RL on verifier scores, and even meta-verifiers) do not. STV roughly doubles accuracy on hard math and lifts it 14x on scientific reasoning tasks (1.5% to 21%). At training time, we additionally train the generator using RL with STV verifier's feedback inside the V-R loop - a procedure we call verifier-in-the-loop training (ViL). Starting from an RL-converged generator, ViL yields a further 33% gain in pass@1. More notably, the generator's standalone pass@1, with no verifier at test time, climbs 30% relative past where standard RL had converged. Hence, the next frontier in reasoning on hard problems may lie in how we train for and with verification.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis