RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling
AuthorsZiyuan Wang, Bohao Tang, Fei Zhang, Shuo Han, Pengfei Liu
Resources
RIBOSPAN is a long-context RNA foundation model designed to understand, represent, generate, and redesign complete transcripts.
Key results
Bidirectional Transformer model capacity
Maximum native pretraining context in nucleotides
RNA sequences in the deduplicated pretraining corpus
Total nucleotide tokens used for pretraining
RIBOSPAN-1K-40 zero-shot mutation-fitness result
RIBOSPAN-10K-15 Pearson correlation for RNA half-life prediction
What the paper found
RIBOSPAN is a 1.61B-parameter RNA foundation model designed to represent complete transcripts at single-nucleotide resolution. It uses a 32-layer bidirectional Transformer with dense self-attention, RoPE positional encoding, single-nucleotide tokenization, and attention-isolated sequence packing, while native masked-language-model pretraining supports contexts up to 10,240 nt. The training corpus combines RNAcentral v26.0, Ensembl 115, and Ensembl Genomes 62, totaling 67.6M RNA sequences and 85.7B nucleotide tokens, with a 15% masking stage followed by 40% continued masking. Compared with short-context models extended by direct RoPE extrapolation or YaRN, native 10K pretraining preserves reconstruction and better balances contextual separation with localized perturbation effects; at 10,240 nt, RIBOSPAN-10K-15 records 0.405962 Additional Context Separation, 0.302668 cross-region same-base similarity, and 0.000785 distal representation diffusion. Frozen evaluations show strong RNA-type organization, especially for long RNAs, while downstream results establish RIBOSPAN as the strongest encoder-only model on RNAGym, with RIBOSPAN-1K-40 reaching 0.2524 Spearman correlation. On mRNABench, RIBOSPAN-10K-15 achieves 0.6522 Pearson correlation for RNA half-life and leads most full-transcript property tasks. The same backbone also supports conditional masked discrete diffusion, AdaLN-Zero conditioning, and synonymous-codon diffusion for joint 5′ UTR, CDS, and 3′ UTR generation, redesign, and protein-preserving optimization.
Original abstract
Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bidirectional RNA foundation model natively pretrained with context lengths up to 10,240 nt. RIBOSPAN combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to enable high-resolution modeling of complete long RNAs. We evaluate the model through nucleotide reconstruction, a controlled long-context representation benchmark, and frozen RNA-type representation analysis. Native 10K pretraining preserves strong reconstruction at 10,240 tokens, while continued pretraining with 40% masking improves recovery under heavy corruption while preserving representation quality. The long-context benchmark further shows that native 10K models maintain strong contextual responsiveness and context-specific representation separation while keeping perturbation-induced representation changes highly localized. Inference-time YaRN scaling recovers much of the contextual organization lost by direct extrapolation of short-context models, but induces substantially greater distal representation diffusion. Frozen-representation evaluations further demonstrate state-of-the-art RNA representation quality, with RIBOSPAN achieving the strongest overall performance across diverse RNA types and retaining a clear advantage on long RNAs. Building on the same backbone, we develop a multidimensionally conditioned discrete-diffusion framework for full-length mRNA generation and redesign, including synonymous-codon diffusion for protein-preserving CDS optimization. Together, RIBOSPAN establishes a powerful long-context foundation for transferable RNA representation learning and full-transcript mRNA design.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.