NTH

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

AuthorsSyeda Anshrah Gillani, Mirza Samad Ahmed Baig

August 22, 2026 3 min read
Watch on YouTube
The one-line take

A large-scale audit finds that LLMs’ doctor recommendations are strongly shaped by ratings, fees, position, and subtle demographic signals that their explanations fail to reveal.

Key results

3024
Randomized choice sets

Frozen conjoint design containing five synthetic physician cards per choice set.

40068
Scored responses

Responses collected across six open-weight models and OpenAI gpt-4o-mini.

31.4
Rating effect

Percentage-point increase from moving the rating from 3.9 to 4.7.

20.0
Fee effect

Percentage-point decrease from raising the visit fee from $90 to $190.

0.03%
Demographic mentions in reasons

Upper bound for gender or ethnicity mentions in stated explanations.

100%
DeepSeek parse failures

deepseek-r1:7b failed the prespecified auditability gate.

What the paper found

This paper audits how large language models recommend physicians using a prespecified randomized choice-based conjoint experiment: 3024 choice sets presented five synthetic family-medicine profiles with independently randomized ratings, review volume and recency, fees, telehealth, experience, affiliation, list position, and names signaling gender and ethnicity. Across 40068 scored responses from six open-weight models and OpenAI’s gpt-4o-mini, average marginal component effects showed that reputation dominated: raising a rating from 3.9 to 4.7 increased recommendation probability by 31.4 percentage points, while increasing the fee from $90 to $190 reduced it by 20.0 percentage points. The audit also rejected demographic parity, finding female-signaled names gained 2.5 percentage points and Hispanic-, South-Asian-, and Black-signaled names gained 2.8, 2.9, and 1.3 percentage points over White-signaled names. First position carried a fee-equivalent advantage of $11 per visit. These effects were largely absent from explanations: demographic attributes appeared in no more than 0.03% of stated reasons, while abstention occurred in only 0.39% of trials. DeepSeek’s deepseek-r1:7b failed the auditability gate with 100% parse failures. An exploratory, non-confirmatory pilot of Anthropic’s Claude-family models suggested similar reputation weighting, but the authors argue that only repeatable behavioral audits—not model self-report—can expose hidden recommendation biases.

Original abstract

Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis