MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
AuthorsHyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
Resources
MedPMC is a large-scale pipeline that turns biomedical papers into cleaner, more useful image-text training data, boosting medical multimodal models and clinical retrieval.
Key results
Source articles processed by the MedPMC curation framework.
Curated pairs produced for multimodal foundation-model training.
Images judged medically relevant in manual review.
Percentage-point average improvement over BMC-CLIP across 26 benchmarks.
Percentage-point gain after replacing the LLaVA-Med vision encoder.
Percentage-point gain on 10,524 Yale New Haven Health System images.
What the paper found
MedPMC, developed by a Yale-led team including researchers affiliated with Microsoft Research, turns permissively licensed PubMed Central literature into continuously updateable multimodal training infrastructure. Its five-stage pipeline uses PubMedBERT for caption-based screening, Vision Transformers for panel detection and medical-image filtering, YOLOv10 for panel separation, and a supervised InternVL-2.5-4B model for caption alignment. Processing 6.1M PMC articles produced 11M medical image-text pairs, with 95.3% of sampled images judged medically relevant, compared with 19.7% in a prior PMC-derived resource. The resulting MedPMC-CLIP, initialized from OpenCLIP ViT-L/14 and trained under the architecture-matched BMC-CLIP setup, improved average zero-shot AUC by 7.1 percentage points across 26 benchmarks spanning 11 specialties, despite using fewer than half as many pairs. Replacing the vision encoder in LLaVA-Med improved medical question answering by 1.9 points on MMMU and 16.9 points on OmniMedVQA; GPT-4 was used to generate synthetic curation annotations, while Microsoft Research contributed affiliated expertise. In a clinical transfer test using 10,524 Yale New Haven Health System dermatology images, MedPMC-CLIP increased morphology-to-image retrieval Recall@5 by 11.7 percentage points. The central result is that high-fidelity filtering, panel decomposition, and image-text alignment can outperform simply scaling noisy biomedical corpora.
Original abstract
Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.