NTH

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

AuthorsJiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

AffiliationsCarnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute · Stanford University, Department of Geophysics · Carnegie Mellon University, Department of Chemistry · Massachusetts Institute of Technology, Plasma Science and Fusion Center · Columbia University, Department of Applied Physics and Applied Mathematics · Princeton University, Department of Astrophysical Sciences · Flatiron Institute, Center for Computational Astrophysics · Engram · Princeton University, Department of Computer Science

October 4, 2026 2 min read
Watch on YouTube
The one-line take

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Key results

26
EurekaBench tasks

Cross-domain scientific mechanism-discovery problems.

306
Scientific insight questions

Expert-verified questions used to assess discoveries.

29.8%
Best AI conditioned score

Claude Fable 5.1’s overall score after scientific-constraint conditioning.

63.7%
Human conditioned score

Reference score from human-discovered mechanisms.

47.4%
GPT 6 Astra predictive accuracy

Near the human predictive-accuracy score of 48.8%.

29.4%
GPT 6 Astra scientific insights

Insight score versus 69.7% for human mechanisms.

What the paper found

EurekaBench evaluates whether AI agents can discover scientific mechanisms rather than merely optimize predictions. It contains 26 expert-designed tasks spanning neuroscience, geophysics, plasma physics, astrophysics, computer science, and chemistry, with 306 verified scientific-insight questions. Agents iteratively query domain simulators, document experiments, submit an interpretable mechanism and executable code, and are scored on scientific constraints, held-out predictive accuracy, and insights derivable from the mechanism; Claude Opus 4.8 and GPT 5.6 Sol serve as agentic judges. The study tests Anthropic models including Claude Fable 5.1 and Claude Opus 5, OpenAI models including GPT 6 Astra and GPT 5.6 Sol, plus DeepSeek V4 Flash and Kimi K3. Claude Fable 5.1 achieves the highest conditioned score among agents at 29.8%, versus 63.7% for human mechanisms. GPT 6 Astra reaches 47.4% predictive accuracy, close to the human 48.8%, but its scientific-insight score is only 29.4%, compared with 69.7% for humans. Across agents, predictive accuracy and insight have limited association, with Kendall’s tau-b of 0.20, showing that fitting observations does not reliably produce understanding. Agents commonly over-optimize numerical accuracy, rely on untested hypotheses, fail to run experiments that reduce uncertainty, and stop after recording informative observations. Directly supplying insight questions raises insight performance, but the authors argue this is artificial; the central deficit is open-ended experimentation and interpretable, extrapolatable mechanism discovery.

Original abstract

When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang

HypoEvolve uses teams of LLM agents and genetic evolution to generate and refine drug-repurposing hypotheses that better align with biological evidence.

Read analysis