NTH

Towards Automating Scientific Review with Google's Paper Assistant Tool

AuthorsRajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, Vincent Cohen-Addad

June 29, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces an AI assistant that helps review scientific papers by checking proofs, validating experiments, and finding errors before human referees do.

Key results

89.7%
SPOT accuracy

PAT detection accuracy on the Math/CS equation-and-proof subset

55.2%
Zero-shot Gemini 3.1 Pro accuracy

Baseline detection accuracy without inference scaling

21.1%
Original SPOT SOTA

Prior state-of-the-art on the same SPOT subset

34%
Relative gain over zero-shot

PAT improvement over zero-shot recall on mathematical errors

4700
Reviewed submissions

Total submissions reviewed across STOC and ICML pilot programs

What the paper found

Google Research’s Paper Assistant Tool (PAT) is presented as an agentic review system built on Gemini Deep Think and designed to automate scientific verification rather than subjective ranking. PAT segments a manuscript into logical units, assigns adaptive compute budgets, runs specialized deep-review agents, and then synthesizes the outputs with Google Search grounding to suppress hallucinations and duplicate critiques. On the SPOT benchmark’s mathematics and computer science equation-and-proof subset, PAT lifts detection accuracy from 55.2% with zero-shot Gemini 3.1 Pro to 89.7%, a 34% relative gain over the zero-shot baseline, while the original SPOT state of the art was 21.1%. In pilot deployments at STOC 2026 and ICML 2026, powered under the hood by an advanced version of Gemini 2.5 Deep Think, PAT reviewed over 4,700 submissions and received strong author feedback: 97% of STOC authors and 92.1% of ICML authors said they would use it again, while 92.7% and 90.7% rated the feedback very or mostly helpful. The paper argues that this kind of AI review could shift peer review from author-facing assistance toward AI-supported reviewer workflows, and eventually toward more automated publication pipelines, while highlighting open problems in hallucination control, parsing, accountability, and adversarial gaming.

Original abstract

Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements, and identifying potential flaws. By utilizing inference scaling techniques, PAT is able to identify deeper issues than a single model call alone, achieving a 34% improvement over zero-shot recall on mathematical errors in the SPOT benchmark. Pilot deployments of PAT as a pre-submission tool for authors at two major Computer Science conferences -- STOC and ICML -- demonstrate its ability to identify critical errors and suggest substantive improvements to research papers. By catching errors early, PAT eases the cognitive burden placed on referees, while preserving their control over the outcomes of the review process.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis