NTH

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

AuthorsRui Meng, Bhavana Dalvi Mishra, Jiefeng Chen, Chun-Liang Li, Palash Goyal, Mihir Parmar, Yiwen Song, Yale Song, Rajarishi Sinha, Parthasarathy Ranganathan, Burak Gokturk, Jinsung Yoon, Tomas Pfister

June 11, 2026 2 min read
Watch on YouTube
The one-line take

ScientistOne is an autonomous research agent that tries to make every claim traceable, aiming to stop hallucinated citations, unverifiable results, and mismatches between code and papers.

Key results

75
ADRS papers

Total audited papers across 5 systems and 5 tasks

337
Bibliography entries

ScientistOne bibliography entries with zero hallucinations

12/12
Score verification

ScientistOne exact score reproduction on ADRS

14/15
Method-code alignment

ScientistOne paper-code consistency on ADRS

98.1%
Numerical claim provenance rate

Native claim provenance pass rate over 639 quantitative claims

1.0600
Parameter Golf score

ScientistOne performance on the live Parameter Golf benchmark

What the paper found

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence, from Google Cloud AI Research, argues that autonomous research systems fail not mainly on score quality but on verifiability, and it introduces Chain-of-Evidence (CoE) so every citation, number, method claim, and conclusion must trace to explicit evidence. The system’s pipeline has three stages: a Problem Investigator that reads up to 100 full-text PDFs per topic, a Parallel Explore-Exploit discovery loop, and a paper writer with a Claim Verifier that checks draft claims against logs, code, and bibliography before finalizing LATEX. On the ADRS benchmark with 75 papers across 5 systems and 5 tasks, ScientistOne is the only system to achieve zero hallucinated references, with 0/337 bibliography entries, perfect score verification at 12/12, and the highest method-code alignment at 14/15; by contrast, baselines show hallucinated reference rates up to 21% and score verification as low as 42%. The paper also reports a native numerical claim provenance rate of 627/639, or 98.1%, and notes that the corrected rate is about 99%. On the solver side, ScientistOne matches or exceeds human experts on all 5 ADRS tasks and reaches 1.0600 on Parameter Golf, where the baseline fails due to a 16MB artifact limit. In broader transfer tests, it earns gold medals on 3D Object Detection and RSNA Brain Tumor, while scaling search width improves TXN by 17% to 4255 and budget scaling improves it by 20% to 4348.

Original abstract

Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis