A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology
AuthorsEugene Vorontsov, Yi Kan Wang, Alican Bozkurt, Adam Casson, Ludmila Tydlitatova, Michal Zelechowski, Ezra E. W. Cohen, Jyoti D. Patel, Max Banaszak, Caitlin McWilliams, Shane Colley, Kate Sasser, Ryan Fukushima, Eric Lefkofsky, Razik Yousfi, Siqi Liu
Resources
This work builds a multimodal oncology foundation model that learns evolving patient representations from clinical records, genomics, and pathology to improve outcome prediction and treatment selection.
Key results
Real-world oncology patients used to develop the oFM
Dimensionality of the longitudinal patient-state representation
oFM AUC compared with 0.563 for curated baseline features
Pooled tAUTOC/SD for oFM embeddings across 11 comparative-treatment cohorts
Comparative-treatment cohorts where oFM benefit ranking exceeded baseline
Total trained parameters in the oFM
What the paper found
The oFM is a longitudinal multimodal oncology foundation model trained by Tempus AI on a 1.67M-patient real-world cohort. It converts daily clinical and molecular events—including diagnoses, treatments, DNA, RNA, and pathology findings—into time-stamped episodes, encodes them with GatorTron-base-2k and Fourier Number Embedding, incorporates H&E whole-slide representations from Virchow2 and PRISM2, and integrates the trajectory with a two-layer temporal Transformer into a 1024-dimensional patient-state embedding. A three-stage curriculum combines TSDAE reconstruction, masked multimodal reconstruction, JEPA-style future latent prediction, VICReg regularization, and Cox survival supervision. Against curated clinical and molecular features, frozen oFM embeddings raised overall-survival AUC to 0.774 from 0.563, with corresponding gains for progression-free survival and treatment response. Across 11 comparative-treatment cohorts, treatment-benefit ranking reached tAUTOC/SD 4.61 versus 1.38 for baseline features, and the oFM ranked benefit better in 9 of 11 cohorts. For interpretation, a sparse autoencoder, latent-space steering, temporal precedence graphs, and retrieval-augmented reasoning with MedGemma 4B connect predictions to clinical and biological mechanisms. The oFM has 421M trained parameters, while its pathology encoders remain frozen. The authors emphasize that these retrospective results support representation quality and hypothesis generation, not validated treatment biomarkers or causal mechanisms.
Original abstract
Precision oncology necessitates a longitudinal model of patient state that captures cancer evolution and treatment over time, integrating multimodal observations. We introduce the oFM, a foundation model developed on a real-world oncology cohort of 1.67 million cancer patients that integrates clinical trajectories with DNA, RNA, and H&E pathology. Patient-level partitions were reserved for training, validation, and testing, with over one million patients used for training. The oFM encodes daily clinical and molecular episodes and, along with pathology images, integrates them over time to produce a patient state embedding. We evaluate frozen oFM embeddings against expert-curated clinical and molecular baseline features. In prognostic benchmarks, the oFM improved AUC for treatment response, progression-free survival, and overall survival (0.774 vs. 0.563 for overall survival). Across 11 comparative-treatment cohorts, the oFM embeddings achieved a three-fold higher pooled and scale-normalized treatment-benefit AUTOC than baseline features with improved benefit ranking in 9 of 11 cohorts, and provided stronger prognostic discrimination within both treatment arms. We also evaluated a mechanism discovery framework that interprets downstream models built on oFM embeddings by linking their predicted outcomes to clinically and biologically grounded mechanisms through an evidence-grounded temporal graph, enabling evaluation in clinical and drug-development applications.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.