One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
AuthorsSebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
Resources
A rigorous audit finds that roughly one in three AI-generated clinical notes contains a verified failure, while showing that the measurement method itself can dramatically change the result.
Key results
Shared consultations from PriMock57, ACI-Bench, and authored scenarios.
Notes generated by three commercial ambient scribes.
Failures surviving adversarial review by Claude Opus 5, GPT-5.5, and GPT-5.4.
177 of 565 notes contained at least one verified failure.
Lenient review verification rate versus 9.3% under the strict instruction.
Allergy-status and medication-list failures.
What the paper found
This paper presents a verified census of three deployed ambient AI clinical scribes, tested on the same 142 consultations from PriMock57, ACI-Bench, and authored scenarios, producing 565 notes without access to patient records. Twelve discovery passes generated 13,678 candidate failures, and an adversarial verification panel using Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.5, with GPT-5.4 resolving disagreements, retained 618 findings. At least one verified failure appeared in 31.3% of notes, concentrated in allergy and medication information, invented patient identity, dropped diagnoses, fabricated dates, and telephone histories rewritten as physical examinations. The taxonomy’s largest cluster contained 111 allergy and medication failures, while a novel unclassified mode recorded a treatment explicitly retracted by the clinician as delivered care. The paper’s central contribution is instrument analysis: on the same evidence, changing Claude Opus 5’s review instruction from adversarially strict to lenient increased candidate verification from 9.3% to 79.0%, while changing reviewer model family had a smaller effect; the full two-family panel altered the strict rate by only one percentage point. Excluding invented identities and dates—errors a patient record might prevent—reduced the note-level rate to 24.8%. The released pipeline includes prompts, model versions, evidence-linked findings, and rerunnable code, arguing that clinical scribe error rates are inseparable from the standard and models used to measure them.
Original abstract
Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.