NTH

SlopShape: Identifying AI-Generated Commercial Web Content

AuthorsJochen Madler

AffiliationsSitefire

October 1, 2026 2 min read
Watch on YouTube
The one-line take

SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.

Key results

2250
Human corpus posts

Pre-ChatGPT commercial blog posts used as human source material

11250
AI mirror posts

Mirrors generated across five AI models

203
Final instrument features

Validated features spanning structural and style dimensions

97.0
Structural detection macro-F1

Human-versus-AI detection on company-disjoint held-out data

96.1
Reworded structural detection macro-F1

Performance after each AI model rewrote its own posts

What the paper found

SlopShape tests whether AI-generated commercial blog posts can be identified from structure rather than wording, extending the StoryScope approach from fiction to business content relevant to Google search and AI-generated answers. The study paired 2,250 pre-ChatGPT human posts from 268 company domains with 11,250 mirrors generated by GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2, and Kimi K2.5. Its commercial-native schema covered 11 dimensions, including argument flow, evidence, audience framing, commercial integration, and voice, producing a 203-feature instrument with 176 structural features after removing style and format artifacts. An XGBoost classifier using structure alone achieved 97.0 macro-F1 on company-disjoint held-out data, compared with 88.1 for style-only features and 98.0 using all features. The structural signal survived self-rewriting: after each model reworded its own posts, performance fell only to 96.1 macro-F1, despite 73% of 13-word sequences being replaced. The recurring AI pattern was a tidy, self-announcing post that states its thesis early, previews its organization, contrasts modern practice with a legacy approach, uses confident institutional voice, and ends with a summary. Structure also supported source attribution: the classifier identified the correct human or AI model for 68.6% of posts, versus a 16.7% six-way chance rate. Human annotations validated the LLM scoring pipeline with human-human kappa of 0.939 and human-model kappa of 0.951, while the authors caution that the results cover single-pass generation and self-rewording, not all humanizer tools or collaborative editing.

Original abstract

Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice. We replicate StoryScope (Russell et al., 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models. A 214-feature instrument, applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.928, human-model 0.946), detects AI posts from its 187 structural features alone at 98.0 macro-F1 on held-out companies, unchanged (98.1) when every AI post is reworded by its own model. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 79.3% are attributed to the correct source against a 16.7% chance rate, and human posts occupy rare structural configurations. All effects replicate StoryScope's, consistent in direction and larger in magnitude. We release pipeline, instrument, prompts, code, and aggregate artifacts.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →