Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice
AuthorsSyeda Anshrah Gillani, Mirza Samad Ahmed Baig
Resources
A large-scale audit finds that LLMs’ doctor recommendations are strongly shaped by ratings, fees, position, and subtle demographic signals that their explanations fail to reveal.
Key results
Frozen conjoint design containing five synthetic physician cards per choice set.
Responses collected across six open-weight models and OpenAI gpt-4o-mini.
Percentage-point increase from moving the rating from 3.9 to 4.7.
Percentage-point decrease from raising the visit fee from $90 to $190.
Upper bound for gender or ethnicity mentions in stated explanations.
deepseek-r1:7b failed the prespecified auditability gate.
What the paper found
This paper audits how large language models recommend physicians using a prespecified randomized choice-based conjoint experiment: 3024 choice sets presented five synthetic family-medicine profiles with independently randomized ratings, review volume and recency, fees, telehealth, experience, affiliation, list position, and names signaling gender and ethnicity. Across 40068 scored responses from six open-weight models and OpenAI’s gpt-4o-mini, average marginal component effects showed that reputation dominated: raising a rating from 3.9 to 4.7 increased recommendation probability by 31.4 percentage points, while increasing the fee from $90 to $190 reduced it by 20.0 percentage points. The audit also rejected demographic parity, finding female-signaled names gained 2.5 percentage points and Hispanic-, South-Asian-, and Black-signaled names gained 2.8, 2.9, and 1.3 percentage points over White-signaled names. First position carried a fee-equivalent advantage of $11 per visit. These effects were largely absent from explanations: demographic attributes appeared in no more than 0.03% of stated reasons, while abstention occurred in only 0.39% of trials. DeepSeek’s deepseek-r1:7b failed the auditability gate with 100% parse failures. An exploratory, non-confirmatory pilot of Anthropic’s Claude-family models suggested similar reputation weighting, but the authors argue that only repeatable behavioral audits—not model self-report—can expose hidden recommendation biases.
Original abstract
Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.