NTH

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

AuthorsBobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

August 20, 2026 3 min read
Watch on YouTube
The one-line take

OmniScientist is an AI researcher that observes raw scientific data across many modalities and disciplines, then turns those observations into experiments and complete research papers.

Key results

36
End-to-end cases

Real-data research cases completed from raw evidence to compiled manuscripts.

6.3
Mean overall score

Mean score on the 7-dimensional review rubric with Claude Sonnet 5.

85%
Perception ablation win rate

Head-to-head judgments won over the blind scalar-feature variant.

21.7%
STEAD noise audit prevalence

Noise-labelled seismic traces containing coherent transient events.

0.851
Radiograph held-out AUC

AUC achieved by mean-entropy plus spatial-patchiness features.

What the paper found

OmniScientist is an end-to-end AI scientist designed to reason over raw scientific evidence rather than text, code, labels, or precomputed scalar features. Its perception layer supports images, signals, audio, video, 3-D structures, trajectories, tables, formulae, sequences, and graphs across 4 evidence families, while 3 autonomous ReAct agents handle ideation, experimentation, and writeup. A deterministic pipeline gates these agents with code-enforced checks for novelty, falsifiability, data leakage, effective sample size, multiple comparisons, execution provenance, anti-HARKing, and numerical claim traceability. Across 36 real-data cases spanning 5 discipline families, the system produced complete compiled manuscripts in every case, achieving a mean overall review score of 6.3 on a 7-dimensional, 0-to-10 rubric with Claude Sonnet 5 as the primary reasoning backbone; comparisons also included OpenAI GPT-5.6, Qwen3.5, Gemma-4, GLM-5.2, and Kimi K2.7. In a paired ablation against a blind interface using scalar features, direct perception won 85% of judgments and improved every evaluation dimension, especially multimodal grounding and scientific significance. The raw-data workflow generated findings such as detecting coherent transients in 21.7% of noise-labelled STEAD seismic traces and identifying a patchiness signature in chest radiographs with a held-out AUC of 0.851. The central result is that lifecycle-wide perception changes the research questions and experiments an AI scientist can formulate, not merely the quality of its final prose.

Original abstract

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis