NTH

Accelerating Scientific Research with Gemini in the Real-World

AuthorsSamuel Schmidgall, Xiaokai Zhu, Marian Shaw, Lin Yang, Valentin Liévin, Jingyun Yang, Yuchen Zhuang, Tim Strother, Alex Bijamov, Min Woo Sun, Anil Palepu, Justin Chen, David Steiner, Jacqueline Shreibati, Wei-Hung Weng, Yilin Zhao, Xingjian Hu, Nicholas Zahn, Sadhya Garg, Julia Kirby, Yuxiang Gan, Jiaoli Li, Divy Thakkar, Shekoofeh Azizi, David Racz, Juraj Gottweis, Vivek Natarajan, Chenglin Wu, Tal Danino, Keran Rong, Haozhe Wang, Benoit Schillings, Yong Cheng, Quoc V. Le, Tao Tu

August 28, 2026 3 min read
Watch on YouTube
The one-line take

A Gemini-based multi-agent scientist conducts experiments, generates hypotheses, and writes research across multiple fields, suggesting a promising but not yet fully validated path toward automated scientific discovery.

Key results

68.0%
MXene-like growth success

Repeated CVD growth success after improved sealing and cleaning protocols.

3 of 4
Biology metric concordance

Morphological metrics whose IPTG-dependent trajectories matched wet-lab measurements.

0.377
HealthBench Hard adjusted score

Agent_H length-adjusted score under the Gemini 3.5 Flash judge.

0.643
HealthBench Professional adjusted score

Agent_H length-adjusted score under the Gemini 3.5 Flash judge.

4%
Severe result hallucinations

Rate after enabling log-based verification and reliability modules.

98.7%
Harmful-prompt refusal rate

Proportion of harmful research directions refused by the safety architecture.

What the paper found

This paper extends Co-Scientist, a Gemini-based multi-agent system, from computational hypothesis generation into execution-grounded scientific research spanning materials science, biology, and computer science. Its pipeline combines evolutionary hypothesis search, Bayesian TrueSkill ranking with Upper Confidence Bound exploration, staged code execution, and log-verified manuscript generation. In materials science, Co-Scientist designed a safer C2Cl6 chemical-vapor-deposition route that produced layered structures resembling Ti3C2Tx MXene, although atomic confirmation remains incomplete; it also achieved first-attempt monolayer growth of MoS2, MoSe2, and WS2, while maintenance changes raised repeated MXene-like growth success to 68.0%. In biology, Gemini 3 Pro Image used leave-one-out interpolation and Best-of-N sampling with N=16 to predict engineered E. coli swarm morphologies across IPTG concentrations, matching wet-lab trajectories on 3 of 4 morphological metrics. In computer science, autonomous architecture search produced Agent_H, an inference-time scaling system using 28–48 candidate responses, clinical auditing, tournament selection, and length control; it reached length-adjusted scores of 0.377 on HealthBench Hard and 0.643 on HealthBench Professional, outperforming single-call GPT-5, OpenAI's GPT-5.6 Sol, Anthropic's Claude Opus 5, and other Gemini baselines on the reported adjusted comparisons. Finally, across 450 blinded expert reviews, log-based hallucination clipping and plagiarism penalties reduced severe result hallucinations to 4%, while the safety system rejected 98.7% of harmful research prompts. The authors emphasize that physical validation, benchmark gaming, residual methodological errors, and human oversight remain essential limitations.

Original abstract

We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypothesis generation, this specialized configuration transitions Co-Scientist into an execution-grounded research partner advancing closed-loop scientific workflows across materials science, biology, and computer science. In materials science, Co-Scientist interfaced with a semi-automated chemical vapor deposition reactor to design a safe precursor route for MXenes; experimental execution produced a lamellar 2D material sharing key structural similarities with the Ti3C2Tx MXene lattice, although further experiments are needed to confirm the atomic structure. Leveraging Gemini 3 Deep Think for rapid, lab-in-the-loop execution, it also tailored growth recipes to laboratory constraints in minutes, enabling single-attempt growth of monolayer MoS2, MoSe2, and WS2 semiconductors. In biology, Co-Scientist predicted emergent swarming phenotypes of engineered E. coli across inducer (IPTG) gradients from sparse imaging data, quantitatively matching unpublished wet-lab morphological measurements. In computer science, Co-Scientist autonomously discovered an inference-time scaling architecture that outperformed six frontier models on HealthBench (Hard and Professional) while reducing potential clinical harm under blinded physician evaluation. Finally, a double-blind study of end-to-end generated papers with 30 domain experts across 450 reviews demonstrates that Co-Scientist's reliability modules reduce hallucination and plagiarism while improving research safety. Together, these results demonstrate progress toward closed-loop multi-agent scientific AI systems capable of accelerating real-world scientific discovery.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis