Language Models Agree With Each Other, Not With Readers
AuthorsKazuki Nakayashiki, Keisuke Watanabe
Resources
Language models often agree with one another far more than they agree with real readers, suggesting that model consensus should not be mistaken for human consensus.
Key results
Naturalistic reader selections used across 120 web documents.
Usable model arms spanning 11 vendors, three countries, generations, sizes, and weight regimes.
Median excess agreement across 153 model pairs.
Excess agreement between two readers on the panel’s 90-document subset.
Excess agreement between frontier models from OpenAI and Anthropic.
Model-versus-human agreement gap after randomizing model selections to match readers’ heavier marking procedure.
What the paper found
Researchers at Glasp evaluate whether language models homogenize against a naturalistic human reference: 2,523 reader mark sets across 120 web documents, created without instructions or payment on a highlighting platform. Their position- and length-controlled excess-overlap estimator compares equal-sized sentence selections after resampling within depth and length bands. Across 18 model arms from 11 vendors, including OpenAI, Anthropic, Google, Meta, Microsoft, NVIDIA, DeepSeek, and others, the median of 153 model pairs reaches +0.093, versus +0.040 for two readers. On a median 70-sentence document, readers each mark 14 sentences and share 4.1, while two models share 8.7; after chance adjustment, the model pair retains 2.8 excess sentences versus 0.6 for readers. GPT-5.4 and Anthropic’s Claude Opus 5 reach +0.203, about twice GPT-4o’s self-agreement of +0.101 and 5.1 times the reader baseline. The result is not explained by determinism, prompt wording, vendor, routing, or generic algorithmic agreement: random-baseline pairs fall within 0.006 of zero, classical extractive pairs have median −0.002, and paraphrasing preserves 91% of self-agreement. Convergence increases with model scale and recency, from +0.056 for two 2024 models to +0.178 for two 2026 models, while agreement with readers rises too but remains within the human range; no model significantly exceeds reader–reader agreement. Matching the readers’ heavier, randomized selection procedure reduces the model gap to +0.082, but it remains above zero. After controlling for sentence depth and length, none of 11 surface features distinguishes model-only from reader-only choices, suggesting models select different content for reasons not visible in sentence form.
Original abstract
Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.