Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark
AuthorsYigeng Jiang, Tengchao Yang, Taoyong Cui, Jiaxing Wan, Yuan Wang, Weida Wang, Zhiyu Liu, Chuyi Peng, Binzhao Luo, Maoli Gao, Huaihai Huang, Yuqianer Zeng, Ziyang Zheng, Dongchen Huang, Chao Chen, Zichao Liu, Weiping Shen, Shuchen Pu, Siyu Zhou, Runmin Ma, Yusong Hu, Fei Chao, Bo Zhang, Xiawu Zheng, Zifu Wang, Lei Bai, Yunqi Cai, Shufei Zhang
This paper introduces a benchmark for AI scientific reasoning in physics and chemistry, along with a multi-agent system that improves performance and lowers inference cost on these hard research tasks.
Key results
Expert-curated benchmark questions
Questions per discipline in PhySciBench
Spearman correlation of composite scoring with blinded expert grading on 196 paired items
Strongest baseline accuracy on PhySciBench
Overall accuracy on PhySciBench
Absolute improvement of DelveAgent over Gemini Deep Research
What the paper found
This paper, from Shanghai AI Lab, Xiamen University, and collaborators including Google DeepMind’s Gemini Deep Research baseline and OpenAI Deep Research, introduces PhySciBench, a 200-question benchmark for physical-science deep research and DelveAgent, a modular multi-agent system built for autonomous scientific workflows. PhySciBench is balanced across 100 physics and 100 chemistry questions and spans six categories: multimodal QA, long-context QA, structured information extraction, scientific reasoning, experimental design, and code generation. The benchmark’s composite scoring pipeline correlates strongly with blinded expert grading at Spearman ρ = 0.80 on 196 paired items, outperforming a single-LLM judge and lexical metrics. On PhySciBench, the strongest baseline, Gemini Deep Research, reaches only 33.5% accuracy, while DelveAgent reaches 41.0%, a gain of 7.5 percentage points, with the largest jumps in structured information extraction and code generation, both up by 15.0 points. The architecture combines an adaptive planning loop, dual-granularity memory, and hierarchical physics-grounded reflection, and ablations show each component matters: removing memory or reflection drops accuracy by 3.0 points, and removing planning drops it by 2.5 points. Across three external benchmarks—HLE, SGI-DR, and FS-Research—DelveAgent averages 23.62% versus Gemini Deep Research’s 20.87%, and in real expert-designed ARPES, Floquet Majorana, and kagome-metal tasks it improves task completion, scientific validity, and hallucination control, showing that domain-specialized orchestration beats scale alone for reliable scientific reasoning.
Original abstract
Deep research agents are Large Language Model (LLM)-based systems designed for autonomous, multi-step scientific reasoning, and they hold immense potential for accelerating research in the physical sciences. However, comprehensive and in-depth evaluations of their capabilities within this domain remain lacking. To address this gap, we introduce PhySciBench, a benchmark highly relevant to physical science research, comprising 200 expert-curated questions, balanced between physics and chemistry, across six task categories that reflect real-world scientific workflows. Evaluations of state-of-the-art models and agent systems on PhySciBench reveal limited performance; even the strongest baseline, Gemini Deep Research, achieves an accuracy of only 33.5%. Analysis of failure cases identifies three recurrent deficiencies: fragility in extended reasoning chains, limited knowledge transfer across steps, and a lack of physics-grounded self-verification. Motivated by these findings, we develop DelveAgent, a modular multi-agent framework equipped with an adaptive planning loop, dual-granularity memory, and a hierarchical physics-grounded reflection mechanism. Across four scientific benchmarks, DelveAgent improves accuracy by up to 7.5 percentage points while reducing inference costs to approximately one-third of the strongest baseline. These results establish the significance of PhySciBench as a critical benchmark for evaluating AI systems in the physical sciences and demonstrate that architectural specialization can effectively enhance the reliability of autonomous scientific research.Our data andcode are publicly available athttps://github.com/yigengjiang/physci-deepresearch.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.