NTH

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it

AuthorsSebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris

September 8, 2026 3 min read
Watch on YouTube
The one-line take

A rigorous audit finds that roughly one in three AI-generated clinical notes contains a verified failure, while showing that the measurement method itself can dramatically change the result.

Key results

142
Consultations audited

Shared consultations from PriMock57, ACI-Bench, and authored scenarios.

565
Notes audited

Notes generated by three commercial ambient scribes.

618
Verified findings

Failures surviving adversarial review by Claude Opus 5, GPT-5.5, and GPT-5.4.

31.3%
Notes with verified failure

177 of 565 notes contained at least one verified failure.

79.0%
Strict-to-lenient verification

Lenient review verification rate versus 9.3% under the strict instruction.

111
Largest failure cluster

Allergy-status and medication-list failures.

What the paper found

This paper presents a verified census of three deployed ambient AI clinical scribes, tested on the same 142 consultations from PriMock57, ACI-Bench, and authored scenarios, producing 565 notes without access to patient records. Twelve discovery passes generated 13,678 candidate failures, and an adversarial verification panel using Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.5, with GPT-5.4 resolving disagreements, retained 618 findings. At least one verified failure appeared in 31.3% of notes, concentrated in allergy and medication information, invented patient identity, dropped diagnoses, fabricated dates, and telephone histories rewritten as physical examinations. The taxonomy’s largest cluster contained 111 allergy and medication failures, while a novel unclassified mode recorded a treatment explicitly retracted by the clinician as delivered care. The paper’s central contribution is instrument analysis: on the same evidence, changing Claude Opus 5’s review instruction from adversarially strict to lenient increased candidate verification from 9.3% to 79.0%, while changing reviewer model family had a smaller effect; the full two-family panel altered the strict rate by only one percentage point. Excluding invented identities and dates—errors a patient record might prevent—reduced the note-level rate to 24.8%. The released pipeline includes prompts, model versions, evidence-linked findings, and rerunnable code, arguing that clinical scribe error rates are inseparable from the standard and models used to measure them.

Original abstract

Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis