NTH

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

AuthorsSeungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal, Ian Wu, Viktor Zaverkin, Spase Petkoski, Daniel R. Schrider, Ilija Dukovski, Francesco Santini, Biljana Mitreska, Yong Jeong, Kyeongha Kwon, Young Min Sim, Dragana Manasova, Arthur Porto, Biljana Mojsoska, Makoto Takamoto, Marko Shuntov, Ruoqi Liu, Hyunjoo Jenny Lee, Niyazi Ulas Dinç, Yehhyun Jo, Sunkyu Han, Chungwoo Lee, Huishan Li, Esther H. R. Tsai, Ergun Simsek, Khushboo Shafi, Yeonseung Chung, Jihye Park, Aleksandar Shulevski, Henrik Christiansen, Yoosang Son, Elly Knight, Amanda Montoya, Jeongyoun Ahn, Christian Langkammer, Heera Moon, Changwon Yoon, Nikola Stikov, Mooseok Jang, Edward Choi, Junhan Kim, Yeon Sik Jung, Woo Youn Kim, Jae Kyoung Kim, Ishraq Md Anjum, Hyun Uk Kim, Drew Bridges, Carolin Lawrence, Xiang Yue, Alice Oh, Akari Asai, Sean Welleck, Graham Neubig

May 27, 2026 3 min read
Watch on YouTube
The one-line take

This study shows that top AI reviewers can already outperform some human reviewers on specific critique quality, but they still have important weaknesses and work best as complements to humans.

Key results

2,960
Annotated criticisms

The study’s atomic unit of analysis is review items, and the paper reports that 45 domain scientists rated 2,960 individual criticisms across the Nature-family corpus.

469 hours
Expert annotation time

The full paper confirms this was the total expert-scientist annotation workload for the item-level meta-reviews.

82
Papers in annotation study

The evaluation was run on 82 Nature-family papers spanning physical, biological, and health sciences.

60.0%
GPT-5.2 fully positive rate

On the composite fully-positive metric, GPT-5.2 exceeds the top-rated human reviewer, as reported in the item-level quality comparison.

48.2%
Top-rated human fully positive rate

This is the human baseline used for the composite comparison against the AI reviewers.

26.0%
AI-only criticism share

The overlap analysis shows AI reviewers surface a distinct set of issues with no human counterpart, and these uncovered AI items remain mostly correct and well-evidenced.

What the paper found

This paper evaluates AI reviewers at the level of atomic review items rather than overall accept/reject scores, using 45 domain scientists who spent 469 hours annotating 2,960 criticisms across 82 Nature-family papers from physics, biology, and health sciences. Each item was judged on correctness, significance, and evidence sufficiency, revealing a nontrivial tradeoff: GPT-5.2 reaches 60.0% fully positive items and outperforms the top-rated human reviewer’s 48.2% on the composite metric (p = 0.009), while Claude Opus 4.5 and Gemini 3.0 Pro remain statistically below or near that benchmark. The distinctive gain is not generic agreement but coverage of different issues: AI reviewers surface 26.0% of criticisms with no human counterpart, and those unique items are still 81.8% correct and 93.5% evidence-sufficient, although AI reviewers overlap with one another far more than humans do, with same-target-same-criticism overlap at 20.9% versus 3.4% for human-human pairs. The main failure modes are also specific: experts identified 16 recurring weaknesses, led by subfield-norm miscalibration, long-context failures that miss information elsewhere in the manuscript or supplement, and over-harsh critiques of minor issues. The authors package the study into PeerReview Bench, a 78-paper benchmark, and CMU Paper Reviewer, an OpenHands-based reviewer service that adds code inspection and citation filtering, but their conclusion is conservative: frontier LLMs can augment human peer review, especially for code-heavy and statistical scrutiny, yet they are not substitutes for a diverse human panel.

Original abstract

With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientists simply view them as probabilistic systems without the expertise to evaluate research, while other researchers are more optimistic about their readiness without concrete evidence. Understanding what AI reviewers do well, where they fall short, and what challenges remain is essential. However, existing evaluations of AI reviewers have focused on whether their verdicts match human verdicts (e.g., score alignment, acceptance prediction), which is insufficient to characterize their capabilities and limits. In this paper, we close this gap through a large-scale expert annotation study, in which 45 domain scientists in Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual criticisms (each targeting one specific aspect of a paper) from human-written and AI-generated reviews of 82 Nature-family papers on correctness, significance, and sufficiency of evidence. On a composite of all three dimensions, a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009), while all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension. AI reviewers' accurate criticisms are also more often rated significant and well-evidenced, and surface a distinct 26% of issues no human raises. However, AI reviewers overlap far more than humans do (21% vs. 3% for cross-reviewer pairs), and exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues. Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis