Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
AuthorsChuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, Yushi Sun
Resources
SDABench tests whether LLMs can genuinely reason about scientific claims, finding that they remain much weaker at inference, causality, and mechanistic explanation than at basic description.
Key results
Real-world scientific data instances in SDABench
Synthetic instances stratified across six capabilities and five domains
Representative open-source, closed-source, and fine-tuned models
Overall accuracy on held-out SDA-Synth tasks
Highest domain-level open-ended accuracy across models
Peak normalized Function Error rate across task types
What the paper found
The paper introduces SDABench, a capability-oriented benchmark that evaluates whether large language models can support scientific claims rather than merely execute analytical workflows. It organizes evaluation around six capabilities—descriptive, exploratory, inferential, predictive, causal, and mechanistic—across Biology, Chemistry, Environment, Geography, and Physics. The benchmark contains 527 real-data instances and 6000 synthetic instances, presented in both multiple-choice and open-ended formats, and evaluates 15 LLMs, including OpenAI’s GPT-5.4, Anthropic’s Claude Sonnet 4.6, Google DeepMind’s Gemini 3.1 Pro, DeepSeek, Qwen, GLM, and Meta’s Llama models. Its construction pipeline combines GPT-4o template generation, semantic causal graphs, programmatic ground-truth derivation, perturbation-based distractors, and human validation. On open-ended synthetic tasks, GPT-5.4 leads with 58.33% overall accuracy, but performance falls sharply beyond descriptive analysis: Environment is the strongest domain at 46.9% overall, while Biology reaches only 22.5% on predictive tasks. A five-stage error taxonomy—Scope, Variable, Function, Relationship, and Conclusion—shows that model scaling reduces early grounding errors, whereas later failures persist: Function Errors peak at 37.7% in predictive tasks, and Relationship Errors dominate mechanistic and exploratory reasoning. The central conclusion is that frontier LLMs can characterize observed data, but they remain unreliable at selecting assumptions, modeling variable relationships, generalizing to unseen cases, and constructing evidence-supported mechanisms.
Original abstract
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.