BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language
AuthorsQizhi Pei, Zhimeng Zhou, Yi Duan, Yiyang Zhao, Wei Li, Han Guo, Liang He, Chengping Li, Chang-Yu Hsieh, Conghui He, Rui Yan, Lijun Wu
Resources
BioMatrix is a single AI system that can read and generate biological sequences, 3D structures, and scientific language for both molecules and proteins.
Key results
Total continual-pretraining tokens across text, molecule, protein, and cross-entity data
Total supervised examples across 80 biological tasks
Tasks where BioMatrix is state-of-the-art or competitive
BioMatrix-4B SMILES exact-match accuracy on text-based molecule generation
BioMatrix-4B error in meV on conditional 3D molecule generation
What the paper found
BioMatrix, built on Qwen3 in 1.7B and 4B sizes, is presented as the first decoder-only biological foundation model to natively unify sequences, 3D structures, and natural language for both molecules and proteins under one next-token objective. Its core technical move is a shared discrete vocabulary that combines SMILES, SELFIES, residue tokens, and learned structure codes from MolStrucTok for molecules and GCP-VQVAE for proteins, eliminating external encoders, projection adapters, and modality-specific heads. The model is continually pretrained on 304.4B tokens spanning text, molecule- and protein-centric multimodal data, plus cross-entity corpora such as BindingDB, CrossDocked2020, and PPIRef, then instruction-tuned on 24,845,248 examples across 80 tasks in 6 categories. Empirically, BioMatrix reports state-of-the-art or competitive results on 77 of 80 tasks, including molecule captioning and text-to-molecule generation, where BioMatrix-4B improves exact-match accuracy from 48.00% with SciReasoner-8B to 65.07%, and property-conditioned conformer generation, where mean absolute error on QM9 electronic targets drops from 205 to 53 meV for HOMO and from 297 to 81 meV for the HOMO-LUMO gap. For proteins, BioMatrix-4B reaches 73.78% total accuracy on MoleculeQA, 90.69% on subcellular localization, 75.50% amino-acid recovery on inverse folding, and 0.963 scTM on unconditional backbone generation. The paper argues that cross-modal gains are largest on tasks that force sequence-structure-language alignment, while residual geometry gaps largely come from finite codebook quantization rather than the autoregressive backbone itself.
Original abstract
We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natural language for both molecules and proteins within a single decoder-only architecture. Existing biological foundation models pursue native multimodality and broad entity coverage separately: those that fuse multiple modalities under a shared objective remain confined to a single entity type, while those spanning multiple entity types either omit explicit structural modeling or rely on adapter-based designs in which the model cannot natively generate the very modalities it can read. BioMatrix closes this gap by mapping molecular sequences (supporting both SMILES and SELFIES notations), molecular structures, protein sequences, protein structures, and natural language into a shared discrete token space through a unified tokenization scheme, so that all modalities are consumed and produced uniformly under a single next-token prediction objective -- without external encoders, projection adapters, or modality-specific output heads. Built upon the Qwen3 language model (1.7B and 4B), BioMatrix is continually pretrained on 304.4 billion tokens spanning general and domain-specific text, sequence and structure views of molecules and proteins, and cross-modal corpora that interleave biomolecular entities with scientific text and link distinct entities through molecule-protein and protein-protein interaction data. After tuning on a comprehensive suite of downstream applications covering 80 tasks across 6 categories -- encompassing single-entity and multi-entity understanding and generation tasks across and within modalities -- BioMatrix achieves state-of-the-art or competitive performance on 77 out of 80 tasks, demonstrating that a single, natively multimodal generalist model can effectively match or surpass specialized approaches across a wide range of biological tasks.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.