NTH

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

AuthorsShiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan, Yujin Wang, Xiandong Li

August 26, 2026 2 min read
Watch on YouTube
The one-line take

G-CARL trains multimodal models to explain medical reports more accurately and empathetically by combining evidence-grounded claim verification with personalized communication checklists.

Key results

2,450
MMedReport benchmark size

Real-world PMRI instances used for evaluation and training.

96.62%
Qwen3-VL-8B-Instruct claim precision

Claim-level precision achieved by G-CARL on MMedReport.

72.18%
Qwen3-VL-8B-Instruct checklist recall

Checklist-level recall achieved by G-CARL on MMedReport.

72.05%
CMB validation QA accuracy

Accuracy achieved without adaptation on the external CMB benchmark.

3.61
CMB professionalism

Open-ended interpretation professionalism score on CMB.

What the paper found

G-CARL introduces Patient-oriented Medical Report Interpretation, or PMRI, an open-ended multimodal task requiring models to explain medical reports in patient-accessible language while considering user queries and dialogue history. Its reinforcement-learning design, built on GRPO, separates three objectives: retrieval-grounded factuality, case-specific communication coverage, and structured reasoning format. The system decomposes responses into atomic claims, retrieves evidence from drug labels, textbooks, and clinical guidelines, then applies dual verification for report relevance and medical support. In parallel, Gemini 3.1 Pro drafts instance-specific checklists that clinicians refine and weight, while GPT-5.2 evaluates the final dimensions of accuracy, satisfaction, and expression. On the 2,450-instance MMedReport benchmark, G-CARL trained with Qwen3-VL-8B-Instruct reaches 96.62% claim precision and 72.18% checklist recall, improving precision by 0.77% and recall by 6.71% over the corresponding base model. On the external CMB benchmark, it achieves 72.05% validation QA accuracy and a 3.61 professionalism score. Ablations show that retrieval, adaptive checklists, and format rewards are complementary; removing retrieval lowers precision, while replacing adaptive checklists with a static rubric reduces overall quality. Clinician pairwise preferences and non-expert comprehension evaluations further favor G-CARL over supervised fine-tuning and holistic MLLM-as-a-Judge rewards, suggesting that evidence-grounded, case-adaptive reinforcement learning can improve both medical reliability and patient-centered communication.

Original abstract

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal generation task that requires models to explain medical reports in accurate and accessible language based on a user's query and dialogue history. These two objectives differ fundamentally in their verifiability, yet remain tightly coupled, making them difficult to optimize jointly under conventional supervised fine-tuning and holistic reinforcement learning paradigms. To address this challenge, we propose G-CARL, a grounded, checklist-aligned reinforcement learning framework that combines multi-source retrieval for atomic claim verification with context-aware, instance-specific weighted checklists for response coverage, providing structured supervision for factuality, user-demand satisfaction, and expression quality without constraining response diversity. We further construct MMedReport, a real-world PMRI benchmark, along with a clinician-designed three-dimensional evaluation protocol. Extensive experiments demonstrate that G-CARL consistently outperforms existing post-training baselines in overall quality, claim-level precision, and checklist recall. Pairwise preference evaluation by clinicians further confirms that G-CARL produces interpretations that are more accurate and better aligned with patient needs.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis