NTH

nnFoundation: 3D Foundation Models for Radiology

AuthorsConstantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, Bálint Kovács, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica Kächele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Krämer, Maximilian Zenk, Tim Rädsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David Kügler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina ...

AffiliationsDivision of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany · Helmholtz Imaging, German Cancer Research Center (DKFZ), Heidelberg, Germany · HIDSS4Health - Helmholtz Information and Data Science School for Health, Helmholtz Association, Karlsruhe/Heidelberg, Germany · Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany · Division of Intelligent Medical Systems, German Cancer Research Center (DKFZ), Heidelberg, Germany · Formerly: Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany · National Center for Tumor Diseases (NCT) Heidelberg, a partnership between the German Cancer Research Center (DKFZ) and Heidelberg University Hospital (UKHD), Heidelberg, Germany · Formerly: Division of Intelligent Medical Systems, German Cancer Research Center (DKFZ), Heidelberg, Germany · German Cancer Consortium (DKTK), DKFZ, core center Heidelberg, Heidelberg, Germany · Medical Faculty Heidelberg, Heidelberg University, Heidelberg, Germany · Helmholtz Metadata Collaboration (HMC) Hub Health, German Cancer Research Center (DKFZ), Heidelberg, Germany · Floy GmbH, Munich, Germany · Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates · Animal Phenotyping Platform, Max-Delbrück-Centrum für Molekulare Medizin in der Helmholtz-Gemeinschaft, Berlin, Germany · AI in Medical Imaging, German Center for Neurodegenerative Diseases (DZNE), Bonn, Germany · Institute of Machine Learning for Biomedical Imaging, Helmholtz Munich, Neuherberg, Germany · School of Computation, Information and Technology, Technical University of Munich, Munich, Germany · Department of Radiation Oncology, Technical University of Munich (TUM), School of Medicine and Health, Klinikum rechts der Isar, Munich, Germany · Digital Technology and Innovation, Siemens Healthineers, Erlangen, Germany · IT Core Facility, German Cancer Research Center (DKFZ), Heidelberg, Germany

September 26, 2026 2 min read
Watch on YouTube
The one-line take

nnFoundation trains complementary 3D convolutional and transformer models on 2.1 million medical scans and shows that matching architecture to the task can improve transfer across diverse radiology applications.

Key results

2.1M
Pretraining volume scale

CT, MRI, and PET volumes used for self-supervised pretraining.

125
Pretraining sources

Institutional and public sources represented in the pretraining corpus.

108
Downstream task count

Tasks spanning segmentation, detection, classification, retrieval, and report generation.

3.0
Segmentation improvement

Dice-point advantage of nnFoundationCNN over the strongest competing model.

99.16%
Low-compute retention

Full-schedule segmentation performance retained using 15% of fine-tuning compute.

8.3
Classification gain

AUROC-point improvement of nnFoundationViT over training from scratch.

What the paper found

nnFoundation introduces two complementary 3D radiology foundation models: nnFoundationCNN for spatially precise segmentation and detection, and nnFoundationViT for globally semantic tasks such as classification, retrieval, and report generation. Using masked autoencoding, the models were pretrained on 2.1M CT, MRI, and PET volumes from 125 sources, then evaluated across 108 downstream tasks and 160,000 evaluation volumes. nnFoundationCNN exceeded the strongest competing model by 3.0 Dice points in segmentation, while nnFoundationViT improved classification over training from scratch by 8.3 AUROC points and achieved a 0.437 F1 score on CT-RATE report generation. A key contribution is dynamic weight adaptation: pretrained kernels and encoder stages are remapped to dataset-specific nnU-Net and nnDetection topologies without additional training, improving transfer by 1.2 Dice points for segmentation and 1.2 mAP points for detection. Pretraining also preserved 99.16% of full-schedule segmentation performance using only 15% of fine-tuning compute, and improved low-data segmentation by 3 Dice points with 10 training cases. The models were integrated into nnU-Net and nnDetection, and frozen visual features were tested with Qwen2.5-VL-3B for report generation. External validation, including workflows involving Siemens Healthineers, supports practical transfer, but the results argue against a single universal architecture: convolutional inductive biases favor localization, while transformers favor global reasoning.

Original abstract

Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report generation, and image retrieval, including evaluations under domain shift, by external partners and in low-data and low-compute regimes. Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging. However, performance follows a consistent task-dependent structure: the convolutional nnFoundation model dominates spatially localized tasks, whereas the transformer-based nnFoundation model excels in tasks requiring global semantic reasoning and in frozen-feature settings. Dynamically aligning the foundation model topology with the dataset characteristics post-hoc further improves transfer across heterogeneous 3D settings. These results show that transferable 3D radiological performance is governed not by a single universal model, but by the interplay of scalable pretraining, complementary architectures, and dataset-aware adaptation. We release nnFoundation models integrated into nnU-Net and nnDetection, enabling immediate application across established radiology workflows.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis