Towards Automating Scientific Review with Google's Paper Assistant Tool
AuthorsRajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, Vincent Cohen-Addad
Resources
This paper introduces an AI assistant that helps review scientific papers by checking proofs, validating experiments, and finding errors before human referees do.
Key results
PAT detection accuracy on the Math/CS equation-and-proof subset
Baseline detection accuracy without inference scaling
Prior state-of-the-art on the same SPOT subset
PAT improvement over zero-shot recall on mathematical errors
Total submissions reviewed across STOC and ICML pilot programs
What the paper found
Google Research’s Paper Assistant Tool (PAT) is presented as an agentic review system built on Gemini Deep Think and designed to automate scientific verification rather than subjective ranking. PAT segments a manuscript into logical units, assigns adaptive compute budgets, runs specialized deep-review agents, and then synthesizes the outputs with Google Search grounding to suppress hallucinations and duplicate critiques. On the SPOT benchmark’s mathematics and computer science equation-and-proof subset, PAT lifts detection accuracy from 55.2% with zero-shot Gemini 3.1 Pro to 89.7%, a 34% relative gain over the zero-shot baseline, while the original SPOT state of the art was 21.1%. In pilot deployments at STOC 2026 and ICML 2026, powered under the hood by an advanced version of Gemini 2.5 Deep Think, PAT reviewed over 4,700 submissions and received strong author feedback: 97% of STOC authors and 92.1% of ICML authors said they would use it again, while 92.7% and 90.7% rated the feedback very or mostly helpful. The paper argues that this kind of AI review could shift peer review from author-facing assistance toward AI-supported reviewer workflows, and eventually toward more automated publication pipelines, while highlighting open problems in hallucination control, parsing, accountability, and adversarial gaming.
Original abstract
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements, and identifying potential flaws. By utilizing inference scaling techniques, PAT is able to identify deeper issues than a single model call alone, achieving a 34% improvement over zero-shot recall on mathematical errors in the SPOT benchmark. Pilot deployments of PAT as a pre-submission tool for authors at two major Computer Science conferences -- STOC and ICML -- demonstrate its ability to identify critical errors and suggest substantive improvements to research papers. By catching errors early, PAT eases the cognitive burden placed on referees, while preserving their control over the outcomes of the review process.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.