NTH

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation

AuthorsShanghua Gao, Ada Fang, Marinka Zitnik

June 21, 2026 2 min read
Watch on YouTube
The one-line take

AutoScientists is a team of AI agents that self-organize to run long scientific experiments, explore multiple hypotheses in parallel, and keep learning from both successes and failures.

Key results

74.40%
BioML-Bench mean percentile

AUTO S CIENTISTS across 24 tasks

66.07%
Autoresearch mean percentile

Matched BioML-Bench baseline

24
BioML-Bench task count

End-to-end biomedical ML tasks

34
GPT target val_bpb

Experiments to reach val_bpb ≈ 0.978 from baseline

0.9730
GPT champion val_bpb

Final best bits-per-byte from a strong champion

217
ProteinGym assay count

Frozen recipe applied across supervised substitution assays

What the paper found

AUTO S CIENTISTS, from Harvard University authors Shanghua Gao, Ada Fang, and Marinka Zitnik, introduces a decentralized multi-agent framework for long-running scientific experimentation that replaces a fixed planner with self-organizing teams, shared experimental state, proposal critique, and dead-end tracking. Built on Anthropic’s Claude Code with Claude Sonnet 4.6 and evaluated on H100 GPUs, the system improved over single-trajectory Autoresearch and other AI agents across BioML-Bench, GPT nanochat optimization, and ProteinGym. On BioML-Bench, it achieved a 74.40% mean leaderboard percentile across 24 tasks versus 66.07% for Autoresearch, with especially strong gains in drug discovery at 64.52%. On GPT training optimization, it reached validation bits-per-byte 0.978 in 34 experiments versus 65 for Autoresearch, and from a strong champion it found 7 accepted improvements to reach 0.9730 while Autoresearch found none in 100 experiments. On ProteinGym, it extended Kermut into a three-GP ensemble that raised ACE2–Spike Spearman correlation from 0.747 to 0.840, and when frozen across all 217 assays it lifted average Spearman’s ρ from 0.657 to 0.700. Ablations show that analyst roles, cross-agent feedback, and self-organization each matter, with the full system consistently outperforming reduced variants under matched compute.

Original abstract

Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existing approaches typically follow a single research trajectory or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration, adapt as experimental evidence changes, or preserve knowledge of failed directions over long-running experiments. We introduce AutoScientists, a decentralized team of AI agents for long-running computational scientific experimentation. Agents interpret a shared experimental state, self-organize into teams around promising hypotheses, critique proposals before using experimental compute, and share successes and failures to reduce redundant exploration. Under matched experimental budgets, AutoScientists improves over prior AI agents across biomedical machine learning, language-model training optimization, and protein fitness prediction. On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest AI agent by +8.33%. On GPT training optimization, AutoScientists reaches a target validation bits-per-byte 1.9x faster than Autoresearch and continues discovering improvements from a starting champion where the single-agent approach finds none (7 vs. 0 accepted improvements). On ProteinGym fitness prediction, AutoScientists discovers a method for ACE2-Spike binding that improves over the current state-of-the-art model by +12.5% in Spearman correlation. Applied without modification across all 217 ProteinGym assays, the same method improves over the prior state of the art by +6.5% (Spearman correlation).

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis