Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA
AuthorsRuinan Jin, Beidi Zhao, Myeongkyun Kang, Qiong Zhang, Xiaoxiao Li
Resources
This paper shows that asking a medical AI to check its own answers can create a dangerous illusion of reliability, where the verifier often just agrees with the original mistake instead of catching it.
Key results
VeriMap is tested on VQA-RAD, SLAKE, PathVQA, PMC-VQA, and MedXpertQA.
The main boundary analysis spans 7 task types × 6 models, producing 42 task-model measurements.
In the logistic mixed-effects analysis, generator error greatly increases the odds that the verifier also errs.
Cross-verification lowers false-positive rate by roughly this amount relative to self-verification.
After four revisions in actor-verifier loops, most initially wrong answers remain locked in by false verification.
What the paper found
Verification Mirage argues that self-verification is not a dependable safety layer for medical vision-language models in visual question answering. The authors introduce VeriMap, a diagnostic framework that separates verifier behavior into discrimination capability and agreement bias, then evaluate six open-weight VLMs—Qwen2.5-VL-7B-Instruct, Gemma-3, Phi-4-Multimodal-Instruct, MedGemma, HuatuoGPT-Vision, and Lingshu—across five datasets: VQA-RAD, SLAKE, PathVQA, PMC-VQA, and MedXpertQA. Across 42 task-model cells spanning seven task types, the dominant regime is the “verification mirage,” where verifier error is high and false acceptance is high simultaneously; on the logistic mixed-effects analysis, a generator error increases verifier-error odds by 57.4× (p<0.001), with coupling varying sharply by model and task. The key novelty is that this failure is task-conditioned: differential diagnosis, causal explanation, and disease classification sit deepest in the mirage, while quantitative measurement is most resistant. Saliency analysis shows a “lazy verifier” effect, with verifiers attending less to image evidence than generators, and cross-verification partially helps, reducing false-positive rate by roughly 12–20 percentage points and verifier error by 2–5 points, but not eliminating the problem. In multi-turn actor-verifier loops, 69.5%–87.1% of initially wrong answers are locked in by false verification after four revisions, showing that reuse amplifies rather than corrects error. The paper’s central claim is precise: in medical VQA, self-verification often reproduces model confidence rather than independently detecting mistakes.
Original abstract
Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a default safety layer for medical visual question answering (VQA). We argue that this practice is fundamentally unreliable. We introduce [METHOD NAME], a diagnostic framework for mapping the reliability boundary of medical VLM self-verification by decomposing verifier behavior into discrimination capability and agreement bias. Because the verifier and answer generator are capacity-coupled, the verifier can overly agree with the generator, creating a verification mirage: a regime with both high verifier error and high agreement bias, driven by false acceptance of incorrect answers. Evaluating six open-weight VLMs across five medical VQA datasets and seven medical tasks, we find that this boundary is strongly task-conditioned. Knowledge-intensive clinical tasks fall deepest into the mirage, simpler tasks are more resistant, and perceptual tasks lie in between. Verification also fails to provide an independent safety signal: logistic mixed-effects analysis shows that verifier error and agreement bias become more likely when the generator is wrong, while saliency analyses show that verifiers under-attend to image evidence relative to generators, a phenomenon we call the lazy verifier. Cross-verification reduces but does not eliminate the mirage. Moreover, when verification is reused in multi-turn actor-verifier loops, most initially wrong answers become locked in by false verification. Since our experiments use clean benchmarks, the observed reliability boundary likely underestimates failures in real clinical deployment.
Read the original paper