NTH

Language Models Agree With Each Other, Not With Readers

AuthorsKazuki Nakayashiki, Keisuke Watanabe

August 6, 2026 3 min read
Watch on YouTube
The one-line take

Language models often agree with one another far more than they agree with real readers, suggesting that model consensus should not be mistaken for human consensus.

Key results

2,523
Reader mark sets

Naturalistic reader selections used across 120 web documents.

18
Model panel

Usable model arms spanning 11 vendors, three countries, generations, sizes, and weight regimes.

0.093
Model-pair median agreement

Median excess agreement across 153 model pairs.

0.040
Human-pair baseline

Excess agreement between two readers on the panel’s 90-document subset.

0.203
GPT-5.4–Claude Opus 5 agreement

Excess agreement between frontier models from OpenAI and Anthropic.

0.082
Reader-procedure-adjusted gap

Model-versus-human agreement gap after randomizing model selections to match readers’ heavier marking procedure.

What the paper found

Researchers at Glasp evaluate whether language models homogenize against a naturalistic human reference: 2,523 reader mark sets across 120 web documents, created without instructions or payment on a highlighting platform. Their position- and length-controlled excess-overlap estimator compares equal-sized sentence selections after resampling within depth and length bands. Across 18 model arms from 11 vendors, including OpenAI, Anthropic, Google, Meta, Microsoft, NVIDIA, DeepSeek, and others, the median of 153 model pairs reaches +0.093, versus +0.040 for two readers. On a median 70-sentence document, readers each mark 14 sentences and share 4.1, while two models share 8.7; after chance adjustment, the model pair retains 2.8 excess sentences versus 0.6 for readers. GPT-5.4 and Anthropic’s Claude Opus 5 reach +0.203, about twice GPT-4o’s self-agreement of +0.101 and 5.1 times the reader baseline. The result is not explained by determinism, prompt wording, vendor, routing, or generic algorithmic agreement: random-baseline pairs fall within 0.006 of zero, classical extractive pairs have median −0.002, and paraphrasing preserves 91% of self-agreement. Convergence increases with model scale and recency, from +0.056 for two 2024 models to +0.178 for two 2026 models, while agreement with readers rises too but remains within the human range; no model significantly exceeds reader–reader agreement. Matching the readers’ heavier, randomized selection procedure reduces the model gap to +0.082, but it remains above zero. After controlling for sentence depth and length, none of 11 surface features distinguishes model-only from reader-only choices, suggesting models select different content for reasons not visible in sentence form.

Original abstract

Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis