Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
AuthorsHongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro, Michael Petrus, Sheikh Momin, Alexandra M. Cardoso Pinto, Rachel Niesen, Laura Sophie Wegner, Dhruv Darji, Jung Moses Koo, Joshua Fieggen, Kapil Narain, Mingde Zeng, Lei Clifton, Linda Shapiro, Fenglin Liu, David A. Clifton
Resources
This paper shows that even strong medical LLMs can be easily tricked by misleading context, and introduces MedMisBench to measure how well they stay correct under pressure.
Key results
retained answer-grounded items in MedMisBench
option-level misleading context-option pairs
mean accuracy on original questions across 11 model configurations
mean post-injection accuracy under focused wrong-option delivery
mean attack success rate under focused wrong-option delivery
reviewed outputs judged wrong with serious potential harm
What the paper found
MedMisBench, from the University of Oxford, University of Washington, University College London, and the University of Waterloo, introduces a new evaluation target for medical LLMs: epistemic resilience, meaning whether a model preserves a correct medical judgment after plausible but false context is injected. The benchmark combines 10,932 answer-grounded medical questions drawn from MedQA, MedMCQA, MedXpertQA, MedJourney, and HLE with 48,889 misleading context-option pairs spanning 5 corruption types and 3 provenance framings. Across 11 model configurations, clean accuracy averaged 71.1%, but focused Type 1 injections dropped performance to 38.0% with 51.5% attack success rate, while Type 2 all-option delivery stayed near 70.5% accuracy yet still produced 18.7% ASR. The most damaging attacks were rule-like fabrications: authority-framed falsehoods reached 69.5% Type 1 ASR, exception poisoning reached 64.1%, and threshold/reference corruption reached 60.9%. In clinician review by a 14-member panel from 7 countries, 38.2% of reviewed responses were worst-case outputs with serious potential harm. On mitigation, Gemini-3.1-pro-preview on HLE saw Type 1 ASR fall from 81.5% to 16.1% with search, while a defensive warning prompt reduced Type 1 ASR by 10.1–14.0 points but did not remove the failure mode.
Original abstract
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: when misleading context is injected into questions that LLMs originally answer correctly, they abandon the correct answer. We call the ability to maintain correct judgment under adversarial context epistemic resilience, and introduce MedMisBench to measure it. MedMisBench contains 10,932 medical question items and 48,889 misleading context-option pairs spanning medical reasoning, agentic capability, and patient-journey evaluation. Across 11 model configurations, mean accuracy falls from 71.1% on original questions to 38.0% under focused misleading context, with 51.5% attack success. The most damaging injections are formal, rule-like fabrications: authority-framed falsehoods reach 69.5% attack success and exception-poisoning claims reach 64.1%. A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases. MedMisBench exposes a structural blind spot in LLM evaluation in medical settings: existing benchmarks measure what models know, but not whether they preserve correct medical judgment under misleading context.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.