Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
AuthorsTaolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao
Resources
SEE tests whether multimodal AI can move beyond explaining science to making reliable, evidence-grounded inferences from real experimental data.
Key results
Multimodal questions covering chemistry, biology, materials science, and 17 subfields
Highest score across 19 evaluated multimodal models
Average SEE accuracy for 15 general-purpose models
Average SEE accuracy for four science-specialized models
Top score using web search and a code interpreter
Explicit recognition of missing visual evidence in text-only ablation cases
What the paper found
Science Edge Evaluation, or SEE, tests whether multimodal large language models can perform evidence-bounded reasoning on authentic laboratory problems rather than merely recall scientific facts. Its benchmark contains 1116 expert-curated questions spanning chemistry, biology, materials science, and 17 fine-grained subfields, using evidence such as NMR and mass spectra, microscopy, Western blots, Cryo-EM, diffraction patterns, and numerical measurements. Across 19 models, including OpenAI’s GPT-5.6-Sol, Google DeepMind’s Gemini 3.1 Pro, Anthropic’s Claude Opus 5, and Alibaba’s Qwen3.8-Max, the best standard accuracy is only 48.7%, while general-purpose models average 34.9% versus 22.2% for science-specialized models. Removing images reveals weak awareness of evidential limits: models explicitly acknowledge missing visual information in just 4.6% of text-only cases. Adding web search and a code interpreter improves the top result to 52.7%, but tool trajectories also create errors through incorrect action selection, visual misinterpretation, and overreliance on retrieved information. SEE’s central finding is that current MLLMs can explain established science yet remain unreliable at extracting precise observations, coordinating evidence across disciplines, maintaining global logical coherence, and limiting conclusions to what experiments actually support. The missing capability for real scientific discovery is therefore not simply more knowledge or tools, but disciplined management of multimodal evidence and uncertainty.
Original abstract
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.