NTH

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

AuthorsSebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi

September 19, 2026 2 min read
Watch on YouTube
The one-line take

ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.

Key results

3199
MIMIC-CXR cases

Cases used in the paired image-swap audit.

44786
Paired trials

All-14 evaluation trials for MedGemma-27B.

4.26%
Flip rate with report

Generated-answer changes when the report was present.

20.94%
Flip rate without report

Generated-answer changes after removing the report.

16.7
Paired image-sensitivity increase

Percentage-point increase without the report.

What the paper found

ModaLens introduces a paired image-swap audit for testing whether report-conditioned medical vision-language models actually use radiographs. In MIMIC-CXR, the researchers evaluated 3,199 cases from 293 patients across 14 questions per case, holding the question and radiology report fixed while replacing each image with another study, usually from the same patient. For MedGemma-27B, the generated answer changed with the substituted image on 4.26% of 44,786 trials when the report was present, versus 20.94% when it was removed, a paired increase of 16.7 percentage points under an explicit yes-or-no instruction. Continuous answer margins showed the same pattern even when the binary answer did not flip: mean absolute margin change rose from 0.690 with the report to 2.307 without it. Attention knockouts suggested that blocking direct access to report tokens from early decoder layers, especially layers 0 through 20, partially restored image sensitivity, while the study cautions that this does not localize all downstream report use. The direction replicated in MedGemma-4B, Qwen3.5-9B, Qwen3.5-27B, and LLaVA-NeXT on Mistral-7B, with experiments run on NVIDIA H100 or H200 GPUs. However, because labels were derived from report text rather than independent image annotations, ModaLens measures image sensitivity and report anchoring—not visual correctness.

Original abstract

A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis