G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
AuthorsShiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan, Yujin Wang, Xiandong Li
Resources
G-CARL trains multimodal models to explain medical reports more accurately and empathetically by combining evidence-grounded claim verification with personalized communication checklists.
Key results
Real-world PMRI instances used for evaluation and training.
Claim-level precision achieved by G-CARL on MMedReport.
Checklist-level recall achieved by G-CARL on MMedReport.
Accuracy achieved without adaptation on the external CMB benchmark.
Open-ended interpretation professionalism score on CMB.
What the paper found
G-CARL introduces Patient-oriented Medical Report Interpretation, or PMRI, an open-ended multimodal task requiring models to explain medical reports in patient-accessible language while considering user queries and dialogue history. Its reinforcement-learning design, built on GRPO, separates three objectives: retrieval-grounded factuality, case-specific communication coverage, and structured reasoning format. The system decomposes responses into atomic claims, retrieves evidence from drug labels, textbooks, and clinical guidelines, then applies dual verification for report relevance and medical support. In parallel, Gemini 3.1 Pro drafts instance-specific checklists that clinicians refine and weight, while GPT-5.2 evaluates the final dimensions of accuracy, satisfaction, and expression. On the 2,450-instance MMedReport benchmark, G-CARL trained with Qwen3-VL-8B-Instruct reaches 96.62% claim precision and 72.18% checklist recall, improving precision by 0.77% and recall by 6.71% over the corresponding base model. On the external CMB benchmark, it achieves 72.05% validation QA accuracy and a 3.61 professionalism score. Ablations show that retrieval, adaptive checklists, and format rewards are complementary; removing retrieval lowers precision, while replacing adaptive checklists with a static rubric reduces overall quality. Clinician pairwise preferences and non-expert comprehension evaluations further favor G-CARL over supervised fine-tuning and holistic MLLM-as-a-Judge rewards, suggesting that evidence-grounded, case-adaptive reinforcement learning can improve both medical reliability and patient-centered communication.
Original abstract
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal generation task that requires models to explain medical reports in accurate and accessible language based on a user's query and dialogue history. These two objectives differ fundamentally in their verifiability, yet remain tightly coupled, making them difficult to optimize jointly under conventional supervised fine-tuning and holistic reinforcement learning paradigms. To address this challenge, we propose G-CARL, a grounded, checklist-aligned reinforcement learning framework that combines multi-source retrieval for atomic claim verification with context-aware, instance-specific weighted checklists for response coverage, providing structured supervision for factuality, user-demand satisfaction, and expression quality without constraining response diversity. We further construct MMedReport, a real-world PMRI benchmark, along with a clinician-designed three-dimensional evaluation protocol. Extensive experiments demonstrate that G-CARL consistently outperforms existing post-training baselines in overall quality, claim-level precision, and checklist recall. Pairwise preference evaluation by clinicians further confirms that G-CARL produces interpretations that are more accurate and better aligned with patient needs.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.