HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
AuthorsJieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang
AffiliationsUniversity of California San Diego Texas A&M University · Carnegie Mellon University Mohamed bin Zayed University of Artificial Intelligence
Resources
HypoEvolve uses teams of LLM agents and genetic evolution to generate and refine drug-repurposing hypotheses that better align with biological evidence.
Key results
Drug-repurposing hypotheses were evaluated across 34 cancer types.
HypoEvolve’s mean external DepMap selectivity score.
HypoEvolve’s mean external Open Targets association score.
The strongest baseline’s DepMap selectivity score.
Final hypotheses produced through crossover or mutation, out of 94 runs.
Standard HypoEvolve computational cost per run.
What the paper found
HypoEvolve frames scientific hypothesis discovery as a generational genetic algorithm for specialized LLM agents. Using OpenAI’s GPT-5.4-mini, one agent generates literature-grounded drug-repurposing hypotheses, another performs pairwise Bradley–Terry comparisons based on cancer specificity, target evidence, and testability, and an evolution agent applies semantic crossover or mutation before a fixed-size population retains the strongest parents and offspring. The system was tested across 34 cancer types using external DepMap CRISPR selectivity and Open Targets association scores, with both measures withheld from the search itself. HypoEvolve reached 0.171 on DepMap and 0.426 on Open Targets, outperforming six baselines; the strongest baseline, Tree of Thoughts, scored 0.115 and 0.329. Its standard configuration used six retained hypotheses, six offspring per generation, and three generations, requiring 206 model calls per run. Evolution produced 87 of 94 final hypotheses through crossover or mutation rather than unchanged copying, while fitness-guided parent selection improved the population’s mean and minimum biological-support scores over random selection. Gains also transferred to held-out cancer types, but the authors emphasize that these are computationally supported proposals, not validated treatments: expert review and experiments remain necessary.
Original abstract
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses. Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work. Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.