NTH

LLM-as-a-Verifier: A General-Purpose Verification Framework

AuthorsJacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu, Yixing Jiang, Chelsea Finn, Marco Pavone, Ion Stoica, Azalia Mirhoseini

July 8, 2026 3 min read
Watch on YouTube
The one-line take

This paper turns LLMs into fine-grained verifiers that can score, rank, and guide agentic tasks more effectively than ordinary judges, improving benchmark performance and even training efficiency.

Key results

86.5%
Terminal-Bench V2 accuracy

state-of-the-art trajectory selection performance

78.2%
SWE-Bench Verified accuracy

state-of-the-art trajectory selection performance

87.4%
RoboRewardBench accuracy

state-of-the-art trajectory preference accuracy

73.3%
MedAgentBench accuracy

state-of-the-art trajectory selection performance

67.13%
PPT full-round-robin accuracy

Probabilistic Pivot Tournament with k=9 on 20-candidate pools

1.8x
LIBERO sample efficiency

improvement over sparse-reward baseline with DSRL-SAC

What the paper found

LLM-as-a-Verifier, from Stanford University, UC Berkeley, and NVIDIA Research, reframes verification as a scaling axis for agentic systems by replacing discrete LLM judge scores with the expectation over scoring-token logits, producing continuous rewards without extra training. Using Gemini 2.5 Flash as the verifier, the framework scales along score granularity, repeated evaluation, and criteria decomposition, raising Terminal-Bench V2 pairwise verification accuracy from 73.1% at G=1 to 77.5% at G=20, from 74.7% at K=1 to 77.5% at K=16, and from 75.2%–76.4% for single criteria to 78.3% when three criteria are ensembled. A cost-efficient Probabilistic Pivot Tournament reduces candidate ranking from O(N^2) to O(Nk) and, on 20-candidate pools, improves selection accuracy from 64.64% for the V1 baseline at 1N budget to 67.13% at 9,630 queried pairs with k=9. In downstream test-time scaling, the verifier sets new state of the art on Terminal-Bench V2 at 86.5%, SWE-Bench Verified at 78.2%, RoboRewardBench at 87.4%, and MedAgentBench at 73.3%. The same fine-grained signal also tracks task progress, reaching Spearman VOC 0.848 on successful Terminal-Bench trajectories and 0.966 on RoboRewardBench, and it improves reinforcement learning sample efficiency, giving about 1.8× higher efficiency on LIBERO with DSRL-SAC and about 1.1× on MATH with GRPO.

Original abstract

Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring additional training. Unlike standard LM judges that prompt LLMs to produce discrete scores for candidate solutions, LLM-as-a-Verifier computes the expectation over the distribution of scoring token logits to generate continuous scores. This probabilistic formulation enables verification to scale along multiple dimensions: (1) score granularity, (2) repeated evaluation, and (3) criteria decomposition. In particular, we show that scaling the scoring granularity leads to better separation between positive and negative solutions, resulting in more calibrated comparisons. Moreover, scaling repeated evaluation and criteria decomposition consistently lead to additional gains in verification accuracy through variance and complexity reduction. We further introduce a cost-efficient ranking algorithm for selecting the best solution among candidates using the verifier's continuous scores. LLM-as-a-Verifier achieves state-of-the-art performance on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%). Beyond verification, the fine-grained signals from LLM-as-a-Verifier can also serve as a proxy for estimating task progress. We build an extension for Claude Code, enabling developers to monitor and improve their own agentic systems. Finally, we show that LLM-as-a-Verifier can provide dense feedback for RL, improving the sample efficiency of SAC and GRPO on robotics and mathematical reasoning benchmarks.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis