NTH

"That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments

AuthorsJason Miklian, John E. Katsos

June 13, 2026 3 min read
Watch on YouTube
The one-line take

This study shows that online accusations of “AI slop” are growing fast, but they often function more as social gatekeeping than as real detection of LLM-generated writing.

Key results

25M
comments analyzed

Total public comments scanned across Hacker News and Reddit

2048
matched controls

Non-accused human comments used in the matched-control test

421
accused parents

Accused parent comments retained for the matched-control analysis

94%
slop share

Share of pejorative mentions using the slop frame by 2026

What the paper found

This study by Jason Miklian and John E. Katsos examines how Reddit and Hacker News users responded to LLM-generated writing by analyzing about 25 million comments from 2023 to 2026, with per-comment classification done using Claude Opus 4.7 from Anthropic. The core finding is that readers did not coordinate on a more accurate detection signal; instead, they rapidly enregistered a pejorative accusation register centered on “AI slop,” whose share of accusation-like comments rose from 2.5 percent to 26.6 percent on Hacker News and from 1.5 percent to 24.4 percent on Reddit, while older inauthenticity terms like “shill” and “astroturf” stayed flat or declined. In a matched-control test of 421 accused parent comments against 2,048 length-matched controls, prose markers that distinguish AI-generated disclosure text from human accusation text—such as lower contraction rate, higher formal adverb density, and greater sentence-length variance—did not predict which human comments were accused, falsifying the idea that these accusations screen for AI at the lay level. The social function instead shifted from mockery toward gatekeeping and structural protest, with “slop” accounting for 94 percent of pejorative mentions by 2026, and the overall sentiment of confirmed accusations becoming more negative over time. The paper argues that this is a case of substitute signaling without screening accuracy: a community-level authenticity police language that survives because it performs boundary maintenance and status signaling, not because it reliably identifies AI-authored text.

Original abstract

Generative AI has made fluent prose cheap to produce, breaking the old promise to readers that good writing meant real thinking. How have readers responded, and what can this tell us about changing anti-AI attitudes? We analyzed 25 million comments from Hacker News and Reddit (2023-2026), combining LLM judgment on 7,500 sampled accusations of AI use, sentiment trajectories, speech-act coding of 300 confirmed accusations of AI use, and a matched-control test of accused versus non-accused parent comments. We found that the pejorative-label share of accusations rose more than tenfold on both platforms while a placebo vocabulary of pre-2022 inauthenticity terms (shill, astroturf) did not. This shift reflected a fast-growing trend of branding any suspicious or seemingly inauthentic prose as "AI slop". The slop frame now constitutes 94 percent of pejorative mentions, with the dominant comments shifting in tone from mockery toward gatekeeping and structural protest. The key surprise comes from a matched-control test which found that prose features that statistically distinguish AI from human text do not predict which human text gets accused as AI. The new accusations work as social gatekeeping of perceived authenticity without actually screening for AI. This research extends signaling theory by showing that substitute signals used socially can grow even when inaccurate if the underlying detection problem cannot be solved at the non-expert level. It shows that AI's effects on writing from the reader side are distinct from those on the production (writer) side. Detection technology cannot resolve this dynamic because the social function of accusations is increasingly to perform social gatekeeping and in-group signaling as opposed to identifying AI-generated writing.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis