Autonomous Scientific Discovery via Iterative Meta-Reflection
AuthorsBingchen Zhao, Sara Beery, Oisin Mac Aodha
Resources
DiscoPER is an LLM-based system that autonomously explores data, tests hypotheses statistically, and reflects on its own findings to uncover new scientific patterns in multimodal datasets.
Key results
Peer-reviewed ecological patterns rediscovered by DiscoPER
Fraction of proposed hypotheses that passed held-out validation
Peer-reviewed ecological patterns rediscovered on the larger benchmark
Fraction of proposed hypotheses that passed held-out validation
DiscoPER without reflection on iNatDisco-800
DiscoPER without reflection on iNatDisco-50K
What the paper found
Autonomous Scientific Discovery via Iterative Meta-Reflection introduces DiscoPER, a code-driven large language model framework that performs open-ended scientific discovery from raw multimodal data without pre-specified research questions. DiscoPER uses a generalized propose–evaluate–reflect loop: it generates executable hypotheses, validates each claim with statistical testing on train and held-out splits, and every five iterations a second-order reflection module analyzes the accumulated accepted and rejected claims to detect gaps, confounds, and compound hypotheses that steer later search. Evaluated on iNatDisco, a new iNaturalist-based ecological benchmark, the system recovers 8 of 9 literature-backed patterns on iNatDisco-800 with a 72.7% hypothesis support rate, and 8 of 12 patterns on iNatDisco-50K with a 74.2% support rate, outperforming classical causal discovery baselines and guided LLM baselines. The paper also shows that removing reflection drops recall to 7 of 9 on iNatDisco-800 and 6 of 12 on iNatDisco-50K, while a counterfactual benchmark confirms DiscoPER follows the modified data rather than memorized ecological priors. Using Claude Sonnet 4.6, Claude Opus 4.6, GPT 5.4, and DeepSeek V4 Pro, the authors show the method can also exploit images through vision-language tools to validate discoveries such as broader mammal geographic ranges and higher-latitude fungal niches.
Original abstract
Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation. However, current systems operate within constrained search spaces or require predefined research questions, limiting their capacity for true open-ended inquiry. Furthermore, while they generate hypotheses iteratively, they largely lack the ability to explicitly synthesize their own accumulated findings to uncover complex, interconnected phenomena. We introduce DiscoPER, an autonomous large language model-powered framework that conducts open-ended research by dynamically generating and executing code to explore datasets without pre-specified research objectives. To ensure rigorous scientific validity, every proposed discovery must pass statistical testing. To overcome the limitations of isolated search, our framework introduces a second-order reasoning mechanism that periodically analyzes its own accumulated discoveries. By treating prior discoveries as empirical data, DiscoPER identifies structural patterns, confounds, and epistemic gaps, actively redirecting hypothesis exploration toward uncharted regions of the search space. The search space is further expanded by incorporating tool use, enabling the system to explore hypotheses beyond structured metadata by seamlessly processing and extracting useful information from multimodal sources like images. Evaluated on iNatDisco, a new multimodal ecological knowledge benchmark with pattern-level ground truth obtained from peer-reviewed literature, DiscoPER recovers 8 of 9 known patterns with a 72.7% hypothesis support rate, outperforming both classical causal discovery and LLM-guided baselines. Ablations show that DiscoPER scales with more data, and confirms the benefits of second-order meta-reflection.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.