Large language models improve physician accuracy but lead to false reliance
AuthorsTirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker
Resources
Citation-backed medical LLMs can substantially improve physician accuracy, but they may also make doctors less likely to challenge confidently supported wrong answers.
Key results
Accuracy before CORA assistance across 46 physicians.
Accuracy after physicians reviewed CORA answers and citations.
Percentage-point improvement from retrieval on contamination-resistant case reports.
Relative AI reliance when physicians initially answered incorrectly and perceived citation support.
Relative self-reliance when CORA was incorrect but its citation appeared supportive.
Relative self-reliance when no citation was perceived as supportive.
What the paper found
The paper introduces CORA, or Citation-Oriented Retrieval Assistant, an agentic retrieval-augmented generation system for dermatology that iteratively reformulates queries, retrieves evidence from guidelines, textbooks, and case reports, reranks passages, and generates answers with linked citations. Across five models—GPT-5, Llama 4, Qwen 2.5, Mistral Large 2, and Gemma 3—CORA preserved baseline accuracy on 2,207 matched DermBenchQA questions and produced especially large gains on the contamination-resistant DermCaseQA set, including a 22.7 percentage-point improvement for Gemma 3. In a within-subject study of 46 physicians making 736 decisions, accuracy rose from 70.8% unaided to 82.6% with CORA, but the citations introduced a grounding-miscalibration risk. When physicians judged that a citation supported a correct CORA answer, adoption of that advice increased to 76.9%, compared with 34% without perceived support. However, when CORA was wrong, perceived citation support reduced resistance to the incorrect recommendation to 34.8%, versus 92.0% without support. The system used Llama-4 Scout for the reader-study generation component, Qwen3-235B-A22B-Instruct-2507 for retrieval orchestration, ChromaDB, Snowflake Arctic Embed V2, and Mixedbread Rerank Large v1. The central finding is that source-linked LLM assistance can improve physician accuracy, while making unsupported or incorrect answers appear sufficiently grounded to weaken clinician scrutiny.
Original abstract
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.