NTH

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

AuthorsZhengrui Guo, Zhengyu Zhang, Jiabo Ma, Yihui Wang, Fengtao Zhou, Yingxue Xu, Ling Liang, Chenglong Zhao, Qi Xie, Jinbang Li, Shujing Guo, Fangyi Han, Zhijian Cen, Ziyi Liu, Cheng Jin, Junlin Hou, Zhixuan Chen, Yu Cai, Lijuan Qu, Shifu Chen, Yueping Liu, Zhe Wang, Xiuming Zhang, Muyan Cai, Li Liang, Hao Chen

May 27, 2026 3 min read
Watch on YouTube
The one-line take

This work introduces a clinically validated lung pathology foundation model that not only performs well across many diagnostic tasks, but also improves pathologist accuracy, speed, and consistency in real-world trials.

Key results

about 2 million
Trainable parameters

PulmoFoundation uses LoRA-based continual pretraining on Virchow2 with only about 2 million trainable parameters, or 0.3% of the 632 million-parameter backbone.

39,503 WSIs / 88 million image patches
Pretraining data

The model was pretrained on 39,503 H&E whole-slide images and 88 million patches from 12 sources.

26,643 WSIs across 32 tasks
Retrospective evaluation scale

Retrospective benchmarking evaluated PulmoFoundation on 26,643 WSIs spanning 32 clinically relevant lung pathology tasks.

1,357 consecutive patients
Prospective cohort

In the registered prospective observational study, 1,357 consecutive patients were enrolled and assessed across 11 diagnostic tasks.

92.3%
Prospective average AUC

Across 11 prospective tasks in routine practice, PulmoFoundation achieved an average AUC of 92.3%.

83.8% to 91.7%
RCT diagnostic accuracy

In the crossover randomized controlled trial with eight pathologists and 4,928 matched reads, AI assistance increased diagnostic accuracy from 83.8% without AI to 91.7% with AI.

What the paper found

PulmoFoundation is a lung-specific foundation model built by continual self-supervised pretraining of Virchow2 with LoRA, adding only about 2 million trainable parameters (0.3% of the 632 million-parameter backbone). It was pretrained on 39,503 H&E whole-slide images and 88 million tiles from 12 sources, then evaluated on 26,643 WSIs across 32 clinical tasks covering biopsy diagnosis, frozen-section triage, resection-based staging and grading, molecular biomarker inference, and survival prediction. The model consistently outperformed pan-cancer baselines including UNI, Virchow2, CHIEF, and GigaPath, with especially large gains on lung-specific morphology tasks such as CK5/6 prediction, P63 prediction, Ki-67 inference, TMB prediction, pleural invasion, and lymph-node metastasis. In retrospective validation, it reached macro AUCs of 0.936 for biopsy tasks, 0.908 for frozen sections, 0.921 for resection diagnosis/staging, and C-indices up to 0.790 for LUAD overall survival. In a prospective observational study of 1,357 consecutive patients, it achieved an average AUC of 92.3% across 11 tasks, including 0.992 for biopsy benign-versus-malignant and 0.970 for frozen-section benign-versus-malignant. Pre-specified triage thresholds reduced second-review burden by 68.8% of biopsy cases and 83.0% of frozen-section cases, while deferring 44.5% of IHC orders at PPVs of 1.000, 0.991, and 0.966. In a crossover randomized controlled trial with eight pathologists and 4,928 matched reads, AI assistance increased diagnostic accuracy from 83.8% to 91.7% (OR=2.23, 95% CI 2.05–2.42), reduced median diagnostic time by 19.6%, and improved Fleiss’ κ from 0.56 to 0.76, showing clinically meaningful workflow benefit rather than benchmark-only performance.

Original abstract

Pathological assessment guides lung cancer diagnosis, treatment selection, and prognostic evaluation, yet current CPath approaches rely on task-specific models for isolated objectives. Although pan-cancer foundation models offer versatility, they lack subspecialty-level depth and have not been evaluated across clinical workflows or prospectively validated in real-world settings. We introduce PulmoFoundation, a multi-center, prospectively validated, randomized controlled trial (RCT)-evaluated foundation model for comprehensive lung pathology assessment across pre-operative, intra-operative, and post-operative care. Built upon Virchow2 via subspecialty-specific pretraining using ~40,000 diagnostic H&E-stained whole-slide images (WSIs), PulmoFoundation was systematically evaluated on ~26,000 WSIs across 32 clinically relevant tasks. In addition to accurately predicting molecular markers and patient survival, our model achieves clinical-grade performance in core diagnostic tasks across biopsy, frozen section, and surgical resection slides. In a registered prospective study of 1,357 patients across 11 diagnostic tasks, our model achieved an average AUC of 92.3%. Using pre-specified triage thresholds, PulmoFoundation could reduce additional second-review burden for 68.8% of biopsies and 83.0% of frozen sections, and defer 44.5% of IHC stain orders, with PPVs of 1.0, 0.991, and 0.966. Beyond prospective validation, we conducted a crossover RCT with eight pathologists, in which AI assistance improved diagnostic accuracy across 4,928 case-reader pairs (91.7% w/ AI vs. 83.8% w/o AI). AI assistance also reduced median diagnostic time by 19.6%, increased diagnostic confidence by 8.7%, and improved inter-rater agreement from moderate (kappa = 0.56) to substantial (kappa = 0.76). Together, these evaluations support PulmoFoundation as a clinically validated decision-support system for lung pathology.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis