NTH

BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

AuthorsQizhi Pei, Zhimeng Zhou, Yi Duan, Yiyang Zhao, Wei Li, Han Guo, Liang He, Chengping Li, Chang-Yu Hsieh, Conghui He, Rui Yan, Lijun Wu

June 25, 2026 3 min read
Watch on YouTube
The one-line take

BioMatrix is a single AI system that can read and generate biological sequences, 3D structures, and scientific language for both molecules and proteins.

Key results

304.4B
Pretraining corpus

Total continual-pretraining tokens across text, molecule, protein, and cross-entity data

24,845,248
Instruction-tuning set

Total supervised examples across 80 biological tasks

77
Task success count

Tasks where BioMatrix is state-of-the-art or competitive

65.07
Molecule text-to-generation EM

BioMatrix-4B SMILES exact-match accuracy on text-based molecule generation

53
QM9 HOMO MAE

BioMatrix-4B error in meV on conditional 3D molecule generation

What the paper found

BioMatrix, built on Qwen3 in 1.7B and 4B sizes, is presented as the first decoder-only biological foundation model to natively unify sequences, 3D structures, and natural language for both molecules and proteins under one next-token objective. Its core technical move is a shared discrete vocabulary that combines SMILES, SELFIES, residue tokens, and learned structure codes from MolStrucTok for molecules and GCP-VQVAE for proteins, eliminating external encoders, projection adapters, and modality-specific heads. The model is continually pretrained on 304.4B tokens spanning text, molecule- and protein-centric multimodal data, plus cross-entity corpora such as BindingDB, CrossDocked2020, and PPIRef, then instruction-tuned on 24,845,248 examples across 80 tasks in 6 categories. Empirically, BioMatrix reports state-of-the-art or competitive results on 77 of 80 tasks, including molecule captioning and text-to-molecule generation, where BioMatrix-4B improves exact-match accuracy from 48.00% with SciReasoner-8B to 65.07%, and property-conditioned conformer generation, where mean absolute error on QM9 electronic targets drops from 205 to 53 meV for HOMO and from 297 to 81 meV for the HOMO-LUMO gap. For proteins, BioMatrix-4B reaches 73.78% total accuracy on MoleculeQA, 90.69% on subcellular localization, 75.50% amino-acid recovery on inverse folding, and 0.963 scTM on unconditional backbone generation. The paper argues that cross-modal gains are largest on tasks that force sequence-structure-language alignment, while residual geometry gaps largely come from finite codebook quantization rather than the autoregressive backbone itself.

Original abstract

We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natural language for both molecules and proteins within a single decoder-only architecture. Existing biological foundation models pursue native multimodality and broad entity coverage separately: those that fuse multiple modalities under a shared objective remain confined to a single entity type, while those spanning multiple entity types either omit explicit structural modeling or rely on adapter-based designs in which the model cannot natively generate the very modalities it can read. BioMatrix closes this gap by mapping molecular sequences (supporting both SMILES and SELFIES notations), molecular structures, protein sequences, protein structures, and natural language into a shared discrete token space through a unified tokenization scheme, so that all modalities are consumed and produced uniformly under a single next-token prediction objective -- without external encoders, projection adapters, or modality-specific output heads. Built upon the Qwen3 language model (1.7B and 4B), BioMatrix is continually pretrained on 304.4 billion tokens spanning general and domain-specific text, sequence and structure views of molecules and proteins, and cross-modal corpora that interleave biomolecular entities with scientific text and link distinct entities through molecule-protein and protein-protein interaction data. After tuning on a comprehensive suite of downstream applications covering 80 tasks across 6 categories -- encompassing single-entity and multi-entity understanding and generation tasks across and within modalities -- BioMatrix achieves state-of-the-art or competitive performance on 77 out of 80 tasks, demonstrating that a single, natively multimodal generalist model can effectively match or surpass specialized approaches across a wide range of biological tasks.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis