AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
AuthorsYunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang, Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See
Resources
AgentIdeaBench tests whether AI systems can actively explore the literature and generate genuinely useful scientific hypotheses, not merely summarize what they have already seen.
Key results
Research subfields spanning five scientific disciplines.
Total models tested, including GPT, Claude, Gemini, Llama, DeepSeek, and Qwen families.
Between-model score variance under Active evaluation relative to Static evaluation.
Ideation-score points gained per year of model knowledge cutoff under Active exploration.
Pearson correlation between Static capability and Active–Static improvement.
Scientific World Modeling improvement on the Qwen3.5 9B backbone.
What the paper found
AgentIdeaBench evaluates scientific ideation as an agentic process rather than a one-shot synthesis task. It covers 100 subfields across computer science, physics, biology, chemistry, and medicine, testing 35 models—including GPT and Claude systems from OpenAI and Anthropic, Google Gemini, Meta Llama, DeepSeek, Qwen, and models relevant to the broader NVIDIA agent ecosystem. In Static mode, models receive five curated papers; in Active mode, they search and fetch literature through Semantic Scholar with a 10-call budget, then produce an identical 80–150-word hypothesis. Three literature-grounded critics score originality, feasibility, clarity, impact, and specificity on a weighted 1–10 scale. Across 28 matched models, Active evaluation produces 4.4× the between-model variance of Static evaluation, exposing capability differences that static tests compress. Performance scales at 1.16 points per year of knowledge cutoff under Active exploration versus 0.54 points per year under Static observation. The Active–Static gain correlates with baseline capability at r=0.69: strong models benefit, while weak models can lose quality. Retrieval primarily improves grounding—feasibility rises by 1.21 points, clarity by 0.63, and specificity by 0.58—while measured originality changes by −0.14 points and is not significant. A replay control shows that the interactive retrieval process, rather than merely receiving better papers, carries most of the gain. An exploratory Scientific World Modeling loop helps Qwen3.5 9B by 0.61 points but fails against compute-matched baselines on frontier DeepSeek models, so its value remains unconfirmed.
Original abstract
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.