Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences
AuthorsMingyang Li, Yurou Liu, Jieping Ye, Bing Su, Ji-Rong Wen, Zheng Wang
Resources
LOGOS is a general-purpose science AI that turns many natural-science problems into one shared language, aiming to make a single foundation model useful across multiple scientific domains.
Key results
Total multitask scientific corpus used for continual pre-training
NatureLM comparison in pocket-conditioned ligand design
Best ligand-design result with full Qwen3 weight initialization
Best pocket-conditioned ligand design score achieved by LOGOS-8B
Best Top-1 accuracy on USPTO-50K
Novel building block rate for LOGOS-8B in unconditional MOF generation
What the paper found
LOGOS, developed by Alibaba Group researchers with Renmin University of China, is a general-purpose generative foundation model for the natural sciences that replaces natural language interfaces with a unified scientific grammar over proteins, antibodies, small molecules, reactions, materials, protein pockets, and protein–ligand complexes. The model tokenizes spatial interactions as sequences and trains autoregressively on 44.87 billion tokens, then post-trains on four downstream generative tasks. A key finding is that more natural-language pretraining tokens hurt scientific performance, while inheriting Qwen3 backbone weights helps; the full weight initialization beats random initialization on pocket-conditioned ligand design, improving Vina from -6.91 to -7.43. Scaling from 1B to 8B parameters consistently improves performance: ligand docking reaches a best Vina of -7.76, retrosynthesis Top-1 rises to 74.8%, MOF generation reaches 45.19% Valid, 39.02% VNU, and 17.78% NBB, and protein editing on GFP reaches 0.93 fitness. LOGOS also generalizes beyond pretraining formats, achieving 85.18% amino-acid recovery on CDR-H1 and 85.93% on CDR-L2, while remaining purely sequence-based and often matching or outperforming domain-specific baselines that require explicit 3D geometry.
Original abstract
In this report, we present LOGOS (Language Of Generative Objects in Science), a scientific generative language model that unifies heterogeneous tasks across the natural sciences within a single autoregressive framework based on a shared scientific grammar. It encodes diverse scientific objects and their spatial interactions as token sequences over a common vocabulary. By representing spatial contact and constraint patterns as discrete tokens, the model captures complex structural interactions in a purely sequential manner, without relying on explicit coordinates or geometric neural networks. This unified representation enables a wide range of downstream tasks to be formulated consistently as next-token prediction in the same grammar space, creating strong alignment between continued multi-domain pre-training and downstream objectives. Across diverse tasks, LOGOS consistently matches or outperforms domain-specific baselines, providing preliminary evidence for the feasibility of "one model fits all" in the natural sciences. We train LOGOS models at different scales (1B, 3B, and 8B parameters) and find a consistent positive correlation between model size and performance. This suggests that the future of AI for Science (AI4S) may not lie in building an independent technical stack that is separated from large language models (LLMs). Instead, it may depend on deeply aligning scientific foundation models with LLMs through shared architectures, shared training paradigms, and shared inference infrastructure, so that LLMs can truly become a new entry point for AI4S. We release the model weights and associated resources to facilitate further research.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.