EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
AuthorsJiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
AffiliationsCarnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute · Stanford University, Department of Geophysics · Carnegie Mellon University, Department of Chemistry · Massachusetts Institute of Technology, Plasma Science and Fusion Center · Columbia University, Department of Applied Physics and Applied Mathematics · Princeton University, Department of Astrophysical Sciences · Flatiron Institute, Center for Computational Astrophysics · Engram · Princeton University, Department of Computer Science
Resources
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.
Key results
Cross-domain scientific mechanism-discovery problems.
Expert-verified questions used to assess discoveries.
Claude Fable 5.1’s overall score after scientific-constraint conditioning.
Reference score from human-discovered mechanisms.
Near the human predictive-accuracy score of 48.8%.
Insight score versus 69.7% for human mechanisms.
What the paper found
EurekaBench evaluates whether AI agents can discover scientific mechanisms rather than merely optimize predictions. It contains 26 expert-designed tasks spanning neuroscience, geophysics, plasma physics, astrophysics, computer science, and chemistry, with 306 verified scientific-insight questions. Agents iteratively query domain simulators, document experiments, submit an interpretable mechanism and executable code, and are scored on scientific constraints, held-out predictive accuracy, and insights derivable from the mechanism; Claude Opus 4.8 and GPT 5.6 Sol serve as agentic judges. The study tests Anthropic models including Claude Fable 5.1 and Claude Opus 5, OpenAI models including GPT 6 Astra and GPT 5.6 Sol, plus DeepSeek V4 Flash and Kimi K3. Claude Fable 5.1 achieves the highest conditioned score among agents at 29.8%, versus 63.7% for human mechanisms. GPT 6 Astra reaches 47.4% predictive accuracy, close to the human 48.8%, but its scientific-insight score is only 29.4%, compared with 69.7% for humans. Across agents, predictive accuracy and insight have limited association, with Kendall’s tau-b of 0.20, showing that fitting observations does not reliably produce understanding. Agents commonly over-optimize numerical accuracy, rely on untested hypotheses, fail to run experiments that reduce uncertainty, and stop after recording informative observations. Directly supplying insight questions raises insight performance, but the authors argue this is artificial; the central deficit is open-ended experimentation and interpretable, extrapolatable mechanism discovery.
Original abstract
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang
HypoEvolve uses teams of LLM agents and genetic evolution to generate and refine drug-repurposing hypotheses that better align with biological evidence.