NTH

Large language models improve physician accuracy but lead to false reliance

AuthorsTirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker

August 13, 2026 2 min read
Watch on YouTube
The one-line take

Citation-backed medical LLMs can substantially improve physician accuracy, but they may also make doctors less likely to challenge confidently supported wrong answers.

Key results

70.8%
Unaided physician accuracy

Accuracy before CORA assistance across 46 physicians.

82.6%
CORA-assisted physician accuracy

Accuracy after physicians reviewed CORA answers and citations.

22.7%
Gemma 3 DermCaseQA gain

Percentage-point improvement from retrieval on contamination-resistant case reports.

76.9%
Adoption of correct advice with support

Relative AI reliance when physicians initially answered incorrectly and perceived citation support.

34.8%
Resistance to incorrect advice with support

Relative self-reliance when CORA was incorrect but its citation appeared supportive.

92.0%
Resistance to incorrect advice without support

Relative self-reliance when no citation was perceived as supportive.

What the paper found

The paper introduces CORA, or Citation-Oriented Retrieval Assistant, an agentic retrieval-augmented generation system for dermatology that iteratively reformulates queries, retrieves evidence from guidelines, textbooks, and case reports, reranks passages, and generates answers with linked citations. Across five models—GPT-5, Llama 4, Qwen 2.5, Mistral Large 2, and Gemma 3—CORA preserved baseline accuracy on 2,207 matched DermBenchQA questions and produced especially large gains on the contamination-resistant DermCaseQA set, including a 22.7 percentage-point improvement for Gemma 3. In a within-subject study of 46 physicians making 736 decisions, accuracy rose from 70.8% unaided to 82.6% with CORA, but the citations introduced a grounding-miscalibration risk. When physicians judged that a citation supported a correct CORA answer, adoption of that advice increased to 76.9%, compared with 34% without perceived support. However, when CORA was wrong, perceived citation support reduced resistance to the incorrect recommendation to 34.8%, versus 92.0% without support. The system used Llama-4 Scout for the reader-study generation component, Qwen3-235B-A22B-Instruct-2507 for retrieval orchestration, ChromaDB, Snowflake Arctic Embed V2, and Mixedbread Rerank Large v1. The central finding is that source-linked LLM assistance can improve physician accuracy, while making unsupported or incorrect answers appear sufficiently grounded to weaken clinician scrutiny.

Original abstract

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis