LLM-as-a-Verifier: A General-Purpose Verification Framework
AuthorsJacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu, Yixing Jiang, Chelsea Finn, Marco Pavone, Ion Stoica, Azalia Mirhoseini
Resources
This paper turns LLMs into fine-grained verifiers that can score, rank, and guide agentic tasks more effectively than ordinary judges, improving benchmark performance and even training efficiency.
Key results
state-of-the-art trajectory selection performance
state-of-the-art trajectory selection performance
state-of-the-art trajectory preference accuracy
state-of-the-art trajectory selection performance
Probabilistic Pivot Tournament with k=9 on 20-candidate pools
improvement over sparse-reward baseline with DSRL-SAC
What the paper found
LLM-as-a-Verifier, from Stanford University, UC Berkeley, and NVIDIA Research, reframes verification as a scaling axis for agentic systems by replacing discrete LLM judge scores with the expectation over scoring-token logits, producing continuous rewards without extra training. Using Gemini 2.5 Flash as the verifier, the framework scales along score granularity, repeated evaluation, and criteria decomposition, raising Terminal-Bench V2 pairwise verification accuracy from 73.1% at G=1 to 77.5% at G=20, from 74.7% at K=1 to 77.5% at K=16, and from 75.2%–76.4% for single criteria to 78.3% when three criteria are ensembled. A cost-efficient Probabilistic Pivot Tournament reduces candidate ranking from O(N^2) to O(Nk) and, on 20-candidate pools, improves selection accuracy from 64.64% for the V1 baseline at 1N budget to 67.13% at 9,630 queried pairs with k=9. In downstream test-time scaling, the verifier sets new state of the art on Terminal-Bench V2 at 86.5%, SWE-Bench Verified at 78.2%, RoboRewardBench at 87.4%, and MedAgentBench at 73.3%. The same fine-grained signal also tracks task progress, reaching Spearman VOC 0.848 on successful Terminal-Bench trajectories and 0.966 on RoboRewardBench, and it improves reinforcement learning sample efficiency, giving about 1.8× higher efficiency on LIBERO with DSRL-SAC and about 1.1× on MATH with GRPO.
Original abstract
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring additional training. Unlike standard LM judges that prompt LLMs to produce discrete scores for candidate solutions, LLM-as-a-Verifier computes the expectation over the distribution of scoring token logits to generate continuous scores. This probabilistic formulation enables verification to scale along multiple dimensions: (1) score granularity, (2) repeated evaluation, and (3) criteria decomposition. In particular, we show that scaling the scoring granularity leads to better separation between positive and negative solutions, resulting in more calibrated comparisons. Moreover, scaling repeated evaluation and criteria decomposition consistently lead to additional gains in verification accuracy through variance and complexity reduction. We further introduce a cost-efficient ranking algorithm for selecting the best solution among candidates using the verifier's continuous scores. LLM-as-a-Verifier achieves state-of-the-art performance on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%). Beyond verification, the fine-grained signals from LLM-as-a-Verifier can also serve as a proxy for estimating task progress. We build an extension for Claude Code, enabling developers to monitor and improve their own agentic systems. Finally, we show that LLM-as-a-Verifier can provide dense feedback for RL, improving the sample efficiency of SAC and GRPO on robotics and mathematical reasoning benchmarks.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.