A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
AuthorsZhengrui Guo, Zhengyu Zhang, Jiabo Ma, Yihui Wang, Fengtao Zhou, Yingxue Xu, Ling Liang, Chenglong Zhao, Qi Xie, Jinbang Li, Shujing Guo, Fangyi Han, Zhijian Cen, Ziyi Liu, Cheng Jin, Junlin Hou, Zhixuan Chen, Yu Cai, Lijuan Qu, Shifu Chen, Yueping Liu, Zhe Wang, Xiuming Zhang, Muyan Cai, Li Liang, Hao Chen
Resources
This work introduces a clinically validated lung pathology foundation model that not only performs well across many diagnostic tasks, but also improves pathologist accuracy, speed, and consistency in real-world trials.
Key results
PulmoFoundation uses LoRA-based continual pretraining on Virchow2 with only about 2 million trainable parameters, or 0.3% of the 632 million-parameter backbone.
The model was pretrained on 39,503 H&E whole-slide images and 88 million patches from 12 sources.
Retrospective benchmarking evaluated PulmoFoundation on 26,643 WSIs spanning 32 clinically relevant lung pathology tasks.
In the registered prospective observational study, 1,357 consecutive patients were enrolled and assessed across 11 diagnostic tasks.
Across 11 prospective tasks in routine practice, PulmoFoundation achieved an average AUC of 92.3%.
In the crossover randomized controlled trial with eight pathologists and 4,928 matched reads, AI assistance increased diagnostic accuracy from 83.8% without AI to 91.7% with AI.
What the paper found
PulmoFoundation is a lung-specific foundation model built by continual self-supervised pretraining of Virchow2 with LoRA, adding only about 2 million trainable parameters (0.3% of the 632 million-parameter backbone). It was pretrained on 39,503 H&E whole-slide images and 88 million tiles from 12 sources, then evaluated on 26,643 WSIs across 32 clinical tasks covering biopsy diagnosis, frozen-section triage, resection-based staging and grading, molecular biomarker inference, and survival prediction. The model consistently outperformed pan-cancer baselines including UNI, Virchow2, CHIEF, and GigaPath, with especially large gains on lung-specific morphology tasks such as CK5/6 prediction, P63 prediction, Ki-67 inference, TMB prediction, pleural invasion, and lymph-node metastasis. In retrospective validation, it reached macro AUCs of 0.936 for biopsy tasks, 0.908 for frozen sections, 0.921 for resection diagnosis/staging, and C-indices up to 0.790 for LUAD overall survival. In a prospective observational study of 1,357 consecutive patients, it achieved an average AUC of 92.3% across 11 tasks, including 0.992 for biopsy benign-versus-malignant and 0.970 for frozen-section benign-versus-malignant. Pre-specified triage thresholds reduced second-review burden by 68.8% of biopsy cases and 83.0% of frozen-section cases, while deferring 44.5% of IHC orders at PPVs of 1.000, 0.991, and 0.966. In a crossover randomized controlled trial with eight pathologists and 4,928 matched reads, AI assistance increased diagnostic accuracy from 83.8% to 91.7% (OR=2.23, 95% CI 2.05–2.42), reduced median diagnostic time by 19.6%, and improved Fleiss’ κ from 0.56 to 0.76, showing clinically meaningful workflow benefit rather than benchmark-only performance.
Original abstract
Pathological assessment guides lung cancer diagnosis, treatment selection, and prognostic evaluation, yet current CPath approaches rely on task-specific models for isolated objectives. Although pan-cancer foundation models offer versatility, they lack subspecialty-level depth and have not been evaluated across clinical workflows or prospectively validated in real-world settings. We introduce PulmoFoundation, a multi-center, prospectively validated, randomized controlled trial (RCT)-evaluated foundation model for comprehensive lung pathology assessment across pre-operative, intra-operative, and post-operative care. Built upon Virchow2 via subspecialty-specific pretraining using ~40,000 diagnostic H&E-stained whole-slide images (WSIs), PulmoFoundation was systematically evaluated on ~26,000 WSIs across 32 clinically relevant tasks. In addition to accurately predicting molecular markers and patient survival, our model achieves clinical-grade performance in core diagnostic tasks across biopsy, frozen section, and surgical resection slides. In a registered prospective study of 1,357 patients across 11 diagnostic tasks, our model achieved an average AUC of 92.3%. Using pre-specified triage thresholds, PulmoFoundation could reduce additional second-review burden for 68.8% of biopsies and 83.0% of frozen sections, and defer 44.5% of IHC stain orders, with PPVs of 1.0, 0.991, and 0.966. Beyond prospective validation, we conducted a crossover RCT with eight pathologists, in which AI assistance improved diagnostic accuracy across 4,928 case-reader pairs (91.7% w/ AI vs. 83.8% w/o AI). AI assistance also reduced median diagnostic time by 19.6%, increased diagnostic confidence by 8.7%, and improved inter-rater agreement from moderate (kappa = 0.56) to substantial (kappa = 0.76). Together, these evaluations support PulmoFoundation as a clinically validated decision-support system for lung pathology.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.