NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
AuthorsYuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng, Bingxiang He, Che Jiang, Junlin Yang, Yuchong Wang, Kaikai Zhao, Weifeng Huang, Kai Tian, Zhenzhao Yuan, Jincheng Zhong, Weizhi Wang, Ning Ding, Bowen Zhou, Kaiyan Zhang
NatureBench tests whether coding agents can actually match published Nature-paper results on real scientific problems, and finds they usually fall short of true scientific discovery.
Key results
NatureBench final benchmark size
Approximate papers crawled from ten Nature-family journals
Methodological translation share among validated successes
Share of failures due to wrong method choice
What the paper found
NatureBench, from Horizon Research, Frontis.AI and Tsinghua University, asks whether coding agents can actually beat the published state of the art on real Nature-family science papers, not just reproduce them. The benchmark contains 90 containerized tasks distilled from about 5,500 papers across ten Nature-family journals and six scientific domains, built through the NatureGym pipeline that filters, verifies, and packages each paper into an isolated environment with hidden ground truth and an information firewall. Under a strict web-search-disabled protocol, ten frontier agents were tested, including Claude Code with Claude Opus 4.6 and 4.7, OpenAI Codex CLI with GPT-5.4 and GPT-5.5, Gemini CLI with Gemini 3.5 Flash, plus DeepSeek-V4-Pro, Qwen 3.7 Max, Kimi K2.6, GLM-5.1, and MiniMax-M2.7. The strongest system, Claude Opus 4.7, surpassed the paper-reported SOTA on only 17.8% of tasks and matched it on 47.8%, while the benchmark’s validity checks and reproduce-mode calibration showed the SOTA anchors were well aligned, with median relative gap near zero on tasks both Claude Opus 4.6 and DeepSeek-V4-Pro could reproduce. The main behavioral finding is that agents usually win by methodological translation, most often converting scientific problems into supervised prediction pipelines, which accounted for 45.5% of validated successes, whereas failures are dominated by wrong method choice at 45.1% and insufficient compute budget at 24.4%.
Original abstract
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction toward discovery on real scientific problems. NatureBench is built on NatureGym, an automated pipeline that constructs a standardized, per-task containerized environment from a source paper, addressing the environment-fragmentation problem that has limited the credibility of prior agent-on-research benchmarks. Evaluating ten frontier agent configurations under a strict web-search-disabled protocol, we find that the strongest model surpasses SOTA on only 17.8% of tasks under the g>0.1 criterion. Analysis of method pathways reveals that agents succeed primarily through methodological translation, converting scientific tasks into familiar supervised prediction problems, rather than through genuine scientific invention. Failures are dominated by wrong method choice and insufficient compute budget, not by task misunderstanding. We release the benchmark, the NatureGym pipeline, and a public leaderboard with maintainer-side reproduction. Code: https://github.com/FrontisAI/NatureBench
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.