ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
AuthorsHangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
Resources
ClinFusion is a vision-focused medical multimodal LLM that combines 2D and 3D image understanding with radiologist-aligned evaluation for more reliable clinical reports and reasoning.
Key results
Total multimodal, medical-text, and general-domain training samples.
ClinFusion-8B score on native 3D abdominal CT multiple-choice VQA.
ClinFusion-8B score under the RoI-grounded report evaluation.
ClinFusion-32B instruction-following score.
Blinded radiologist evaluation cases spanning CT and X-ray.
Highest reported correlation between automatic evaluation and expert rankings.
What the paper found
ClinFusion, developed by Alibaba Group’s DAMO Academy, is a vision-centric multimodal large language model system built on Qwen3-VL that targets the central difficulty of medical AI: integrating fine-grained 2D images with native 3D CT and MRI volumes. Its compositional vision encoder combines Qwen ViT with DINOv2, ConvNeXt, and a dedicated PE-3D encoder through Cascade Spatial-Aware Locality, or CaSL, Fusion; localized cross-attention progressively enriches aligned visual tokens, while 2D anchor features guide depth-aware 3D fusion. Trained on 22.2M samples, ClinFusion-8B scores 80.2 on AMOS-MCQ and achieves a 37.8 F1 on CheXpert-Plus report generation, outperforming Hulu-Med-7B and competing strongly with proprietary systems including OpenAI’s GPT-5.2 and Google’s Gemini-3-Flash. The paper also introduces MedIF-Bench, which measures strict clinical instruction compliance, and an RoI-grounded report metric that uses clinical indication, anatomical focus, and an LLM-as-a-judge—OpenAI GPT-4.1—to classify matched, missed, and hallucinated findings. ClinFusion-32B reaches 98.9 Overall-IF, while agentic extensions add retrieval-augmented generation and specialist tools for segmentation and disease assessment. In a blinded study spanning 300 CT and X-ray cases, six board-certified radiologists ranked the tool-augmented system highest; the RoI metric correlated most strongly with expert rankings, reaching Kendall’s tau of 0.511. Across the reported suite, ClinFusion surpasses leading open-source medical models on 20 of 24 benchmarks and outperforms GPT-5.2 and Gemini-3-Flash on 13 of 16 multimodal benchmarks.
Original abstract
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.