ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
AuthorsWanghan Xu, Shuo Li, Tianlin Ye, Qinglong Cao, Yixin Chen, Hengjian Gao, Yiheng Wang, Qi Li, Kun Li, Sheng Xu, Shengdu Chai, Fangchen Yu, Xiangyu Zhao, Zhangrui Zhao, Weijie Ma, Zijie Guo, Haoyu Zhou, Haoxiang Yin, Lixue Cheng, Chaofan Hu, Haoxuan Li, Lu Mi, Xuxuan Xie, Yifan Zhou, Ruizhe Chen, Zhiwang Zhou, Xingjian Guo, Yuhao Zhou, Xuming He, Shengyuan Xu, Xinyu Gu, Jiamin Wu, Mianxin Liu, Chunfeng Song, Fenghua Ling, Dongzhan Zhou, Shixiang Tang, Yuqiang Li, Mao Su, Peng Ye, Siqi Sun, Bin Wang, Xue Yang, Zhenfei Yin, Tianfan Fu, Guangtao Zhai, Wanli Ouyang, Bo Zhang, Lei Bai, Wenlong Zhang
Resources
ResearchClawBench is a benchmark that tests whether AI agents can truly rediscover real scientific papers end to end, revealing that today’s best systems still fall far short.
Key results
real paper-derived tasks in the benchmark
scientific domains covered
best autonomous agent average score
best ResearchHarness LLM average score
mean score across native LLM baselines
best-per-task autonomous-agent frontier mean
What the paper found
ResearchClawBench, from Shanghai Artificial Intelligence Laboratory, is a benchmark for end-to-end autonomous scientific research that converts 40 real papers into executable tasks across 10 domains, including Astronomy, Chemistry, Earth Science, Energy, Information Science, Life Science, Materials, Mathematics, Neuroscience, and Physics. Each task hides the target paper and evaluates systems with expert-curated multimodal rubrics over the full research loop: reading literature, analyzing raw data, running code, and assembling a final report. The benchmark sets 50 points as the target-paper re-discovery boundary, so scores above 50 indicate discovery beyond the source paper. Under a unified protocol, the strongest autonomous agent, Claude Code, averages 21.5, while the best ResearchHarness LLM, Claude-Opus-4.7, averages 20.7; the LLM frontier mean is 26.5, and even the best autonomous-agent frontier mean is only 24.6. Error analysis over 280 runs shows failures concentrate in experiment-design mismatch, evidence mismatch, and missing scientific core, rather than simple execution failure. The paper also situates ResearchClawBench against systems and products from Anthropic, OpenAI, DeepSeek, Google, xAI, Qwen, Kimi, MiMo, and others, showing that current frontier models and agent scaffolds still fall far short of reliable target-paper-level scientific re-discovery.
Original abstract
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific research across 40 tasks from 10 scientific domains. Each task is grounded in a real published paper, provides related literature and raw data, and hides the target paper during evaluation. Expert-curated multimodal rubrics decompose the target scientific artifacts into weighted criteria, enabling evaluation of target-paper-level re-discovery while leaving room for new discovery. We evaluate seven autonomous research (auto-research) agents under a unified protocol and seventeen native LLMs through the lightweight ResearchHarness. Current systems remain far from reliable re-discovery: the strongest autonomous agent, Claude Code, averages 21.5, and the strongest ResearchHarness LLM, Claude-Opus-4.7, averages 20.7, with an LLM frontier mean of only 26.5. Error analysis shows that failures concentrate in experimental protocol mismatch, evidence mismatch, and missing scientific core. ResearchClawBench provides a reproducible evaluation frontier for measuring progress toward autonomous scientific research.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.