On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
AuthorsSeungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal, Ian Wu, Viktor Zaverkin, Spase Petkoski, Daniel R. Schrider, Ilija Dukovski, Francesco Santini, Biljana Mitreska, Yong Jeong, Kyeongha Kwon, Young Min Sim, Dragana Manasova, Arthur Porto, Biljana Mojsoska, Makoto Takamoto, Marko Shuntov, Ruoqi Liu, Hyunjoo Jenny Lee, Niyazi Ulas Dinç, Yehhyun Jo, Sunkyu Han, Chungwoo Lee, Huishan Li, Esther H. R. Tsai, Ergun Simsek, Khushboo Shafi, Yeonseung Chung, Jihye Park, Aleksandar Shulevski, Henrik Christiansen, Yoosang Son, Elly Knight, Amanda Montoya, Jeongyoun Ahn, Christian Langkammer, Heera Moon, Changwon Yoon, Nikola Stikov, Mooseok Jang, Edward Choi, Junhan Kim, Yeon Sik Jung, Woo Youn Kim, Jae Kyoung Kim, Ishraq Md Anjum, Hyun Uk Kim, Drew Bridges, Carolin Lawrence, Xiang Yue, Alice Oh, Akari Asai, Sean Welleck, Graham Neubig
Resources
This study shows that top AI reviewers can already outperform some human reviewers on specific critique quality, but they still have important weaknesses and work best as complements to humans.
Key results
The study’s atomic unit of analysis is review items, and the paper reports that 45 domain scientists rated 2,960 individual criticisms across the Nature-family corpus.
The full paper confirms this was the total expert-scientist annotation workload for the item-level meta-reviews.
The evaluation was run on 82 Nature-family papers spanning physical, biological, and health sciences.
On the composite fully-positive metric, GPT-5.2 exceeds the top-rated human reviewer, as reported in the item-level quality comparison.
This is the human baseline used for the composite comparison against the AI reviewers.
The overlap analysis shows AI reviewers surface a distinct set of issues with no human counterpart, and these uncovered AI items remain mostly correct and well-evidenced.
What the paper found
This paper evaluates AI reviewers at the level of atomic review items rather than overall accept/reject scores, using 45 domain scientists who spent 469 hours annotating 2,960 criticisms across 82 Nature-family papers from physics, biology, and health sciences. Each item was judged on correctness, significance, and evidence sufficiency, revealing a nontrivial tradeoff: GPT-5.2 reaches 60.0% fully positive items and outperforms the top-rated human reviewer’s 48.2% on the composite metric (p = 0.009), while Claude Opus 4.5 and Gemini 3.0 Pro remain statistically below or near that benchmark. The distinctive gain is not generic agreement but coverage of different issues: AI reviewers surface 26.0% of criticisms with no human counterpart, and those unique items are still 81.8% correct and 93.5% evidence-sufficient, although AI reviewers overlap with one another far more than humans do, with same-target-same-criticism overlap at 20.9% versus 3.4% for human-human pairs. The main failure modes are also specific: experts identified 16 recurring weaknesses, led by subfield-norm miscalibration, long-context failures that miss information elsewhere in the manuscript or supplement, and over-harsh critiques of minor issues. The authors package the study into PeerReview Bench, a 78-paper benchmark, and CMU Paper Reviewer, an OpenHands-based reviewer service that adds code inspection and citation filtering, but their conclusion is conservative: frontier LLMs can augment human peer review, especially for code-heavy and statistical scrutiny, yet they are not substitutes for a diverse human panel.
Original abstract
With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientists simply view them as probabilistic systems without the expertise to evaluate research, while other researchers are more optimistic about their readiness without concrete evidence. Understanding what AI reviewers do well, where they fall short, and what challenges remain is essential. However, existing evaluations of AI reviewers have focused on whether their verdicts match human verdicts (e.g., score alignment, acceptance prediction), which is insufficient to characterize their capabilities and limits. In this paper, we close this gap through a large-scale expert annotation study, in which 45 domain scientists in Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual criticisms (each targeting one specific aspect of a paper) from human-written and AI-generated reviews of 82 Nature-family papers on correctness, significance, and sufficiency of evidence. On a composite of all three dimensions, a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009), while all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension. AI reviewers' accurate criticisms are also more often rated significant and well-evidenced, and surface a distinct 26% of issues no human raises. However, AI reviewers overlap far more than humans do (21% vs. 3% for cross-reviewer pairs), and exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues. Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.