NTH

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

AuthorsJieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang

AffiliationsUniversity of California San Diego Texas A&M University · Carnegie Mellon University Mohamed bin Zayed University of Artificial Intelligence

September 22, 2026 2 min read
Watch on YouTube
The one-line take

HypoEvolve uses teams of LLM agents and genetic evolution to generate and refine drug-repurposing hypotheses that better align with biological evidence.

Key results

34
Cancer types evaluated

Drug-repurposing hypotheses were evaluated across 34 cancer types.

0.171
DepMap selectivity

HypoEvolve’s mean external DepMap selectivity score.

0.426
Open Targets association

HypoEvolve’s mean external Open Targets association score.

0.115
Tree of Thoughts DepMap baseline

The strongest baseline’s DepMap selectivity score.

87
Final hypotheses from variation

Final hypotheses produced through crossover or mutation, out of 94 runs.

206
Model calls per run

Standard HypoEvolve computational cost per run.

What the paper found

HypoEvolve frames scientific hypothesis discovery as a generational genetic algorithm for specialized LLM agents. Using OpenAI’s GPT-5.4-mini, one agent generates literature-grounded drug-repurposing hypotheses, another performs pairwise Bradley–Terry comparisons based on cancer specificity, target evidence, and testability, and an evolution agent applies semantic crossover or mutation before a fixed-size population retains the strongest parents and offspring. The system was tested across 34 cancer types using external DepMap CRISPR selectivity and Open Targets association scores, with both measures withheld from the search itself. HypoEvolve reached 0.171 on DepMap and 0.426 on Open Targets, outperforming six baselines; the strongest baseline, Tree of Thoughts, scored 0.115 and 0.329. Its standard configuration used six retained hypotheses, six offspring per generation, and three generations, requiring 206 model calls per run. Evolution produced 87 of 94 final hypotheses through crossover or mutation rather than unchanged copying, while fitness-guided parent selection improved the population’s mean and minimum biological-support scores over random selection. Gains also transferred to held-out cancer types, but the authors emphasize that these are computationally supported proposals, not validated treatments: expert review and experiments remain necessary.

Original abstract

Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses. Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work. Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis