ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
AuthorsRui Meng, Bhavana Dalvi Mishra, Jiefeng Chen, Chun-Liang Li, Palash Goyal, Mihir Parmar, Yiwen Song, Yale Song, Rajarishi Sinha, Parthasarathy Ranganathan, Burak Gokturk, Jinsung Yoon, Tomas Pfister
Resources
ScientistOne is an autonomous research agent that tries to make every claim traceable, aiming to stop hallucinated citations, unverifiable results, and mismatches between code and papers.
Key results
Total audited papers across 5 systems and 5 tasks
ScientistOne bibliography entries with zero hallucinations
ScientistOne exact score reproduction on ADRS
ScientistOne paper-code consistency on ADRS
Native claim provenance pass rate over 639 quantitative claims
ScientistOne performance on the live Parameter Golf benchmark
What the paper found
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence, from Google Cloud AI Research, argues that autonomous research systems fail not mainly on score quality but on verifiability, and it introduces Chain-of-Evidence (CoE) so every citation, number, method claim, and conclusion must trace to explicit evidence. The system’s pipeline has three stages: a Problem Investigator that reads up to 100 full-text PDFs per topic, a Parallel Explore-Exploit discovery loop, and a paper writer with a Claim Verifier that checks draft claims against logs, code, and bibliography before finalizing LATEX. On the ADRS benchmark with 75 papers across 5 systems and 5 tasks, ScientistOne is the only system to achieve zero hallucinated references, with 0/337 bibliography entries, perfect score verification at 12/12, and the highest method-code alignment at 14/15; by contrast, baselines show hallucinated reference rates up to 21% and score verification as low as 42%. The paper also reports a native numerical claim provenance rate of 627/639, or 98.1%, and notes that the corrected rate is about 99%. On the solver side, ScientistOne matches or exceeds human experts on all 5 ADRS tasks and reaches 1.0600 on Parameter Golf, where the baseline fails due to a 16MB artifact limit. In broader transfer tests, it earns gold medals on 3D Object Detection and RSNA Brain Tumor, while scaling search width improves TXN by 17% to 4255 and budget scaling improves it by 20% to 4348.
Original abstract
Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.