NTH

AI systems out-persuade expert humans

AuthorsKobi Hackenburg, Caroline Wagner, Luke Hewitt, Ben M. Tappin, Ed Saunders, Hannah Rose Kirk, Helen Margetts, Christopher Summerfield

June 25, 2026 3 min read
Watch on YouTube
The one-line take

This study shows that AI can already out-persuade even expert humans in conversation, raising major questions about how persuasive systems could be used in politics, fundraising, and everyday influence.

Key results

18978
total conversations

Across the four preregistered experiments

6923
total people

Total participants involved in the experiments

8.2%
layperson advantage

AI versus random laypeople on attitude persuasion

5.6%
selected layperson advantage

AI versus tournament-selected laypeople on attitude persuasion

4.6%
elite debater advantage

AI versus elite debaters on attitude persuasion

10.8%
Save the Children donation advantage

Claude Opus 4.6 versus professional canvassers on real-money giving

What the paper found

This preregistered Oxford and UK AI Security Institute study shows that frontier conversational AI can out-persuade elite humans in live text debates. Across 18,978 conversations from 6,923 people, AI systems including Claude Opus 4.1 and 4.6, ChatGPT-4o, GPT-5.4, Grok 4.20, and Gemini 2.5 Pro beat random laypeople by 8.2 percentage points, tournament-selected laypeople by 5.6 points, elite debaters by 4.6 points, and professional canvassers by 5.9 points on post-conversation attitude change. In Study 2, coaching elite debaters increased their output by 19% more words and 54% more fact-checkable claims, but only raised persuasiveness by 1.0 point; the gap vanished only when AI was throttled to human throughput, with constrained AI limited to 51 words and 92 seconds per reply matching elite debaters. Mechanistically, persuasive impact tracked fact density tightly: unconstrained AI averaged about 37 fact-checkable claims per conversation versus about 12 under the constraint, and the overall relationship between facts and impact reached R2 = 0.89. The advantage also generalized beyond attitudes: in a real-money donation task to Save the Children, Claude Opus 4.6 beat professional canvassers by 10.8 points of a £1 bonus, raising both the share who donated and the amount donated among donors. The paper’s core claim is that AI’s edge comes less from charisma than from faster, denser information delivery.

Original abstract

Many societal decisions are settled by contests of persuasion. Conversational AI is a powerful new entrant in these contests, but whether it can out-persuade skilled and highly incentivized humans has remained unclear. Here, in a series of four preregistered experiments (n = 18,978 conversations from 6,923 people), we pitted AI systems against a range of human persuaders, including laypeople, winners of a separately preregistered four-round online persuasion tournament, professional canvassers, and world championship debaters. We found that AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses. In a follow-up study, AI's advantage persisted after experts received a coaching tool that let them practice against the AI that beat them, review their performance history, and see what AI would have said at key moments. We found converging evidence that AI's advantage stemmed from rapidly deploying larger quantities of information: after coaching, expert humans could tie an AI constrained to respond at human speeds and with human-length messages. In a final study, we show that AI's advantage extends to consequential real-world behavior: AI was nearly 3x more effective than professional canvassers from a UK fundraising firm at raising real-money donations to Save the Children. Together, these results establish that frontier AI systems out-persuade expert humans in conversation, with significant implications for political communication.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis