ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
AuthorsSebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.
Key results
Cases used in the paired image-swap audit.
All-14 evaluation trials for MedGemma-27B.
Generated-answer changes when the report was present.
Generated-answer changes after removing the report.
Percentage-point increase without the report.
What the paper found
ModaLens introduces a paired image-swap audit for testing whether report-conditioned medical vision-language models actually use radiographs. In MIMIC-CXR, the researchers evaluated 3,199 cases from 293 patients across 14 questions per case, holding the question and radiology report fixed while replacing each image with another study, usually from the same patient. For MedGemma-27B, the generated answer changed with the substituted image on 4.26% of 44,786 trials when the report was present, versus 20.94% when it was removed, a paired increase of 16.7 percentage points under an explicit yes-or-no instruction. Continuous answer margins showed the same pattern even when the binary answer did not flip: mean absolute margin change rose from 0.690 with the report to 2.307 without it. Attention knockouts suggested that blocking direct access to report tokens from early decoder layers, especially layers 0 through 20, partially restored image sensitivity, while the study cautions that this does not localize all downstream report use. The direction replicated in MedGemma-4B, Qwen3.5-9B, Qwen3.5-27B, and LLaVA-NeXT on Mistral-7B, with experiments run on NVIDIA H100 or H200 GPUs. However, because labels were derived from report text rather than independent image annotations, ModaLens measures image sensitivity and report anchoring—not visual correctness.
Original abstract
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.