SlopShape: Identifying AI-Generated Commercial Web Content
AuthorsJochen Madler
AffiliationsSitefire
Resources
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.
Key results
Pre-ChatGPT commercial blog posts used as human source material
Mirrors generated across five AI models
Validated features spanning structural and style dimensions
Human-versus-AI detection on company-disjoint held-out data
Performance after each AI model rewrote its own posts
What the paper found
SlopShape tests whether AI-generated commercial blog posts can be identified from structure rather than wording, extending the StoryScope approach from fiction to business content relevant to Google search and AI-generated answers. The study paired 2,250 pre-ChatGPT human posts from 268 company domains with 11,250 mirrors generated by GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek V3.2, and Kimi K2.5. Its commercial-native schema covered 11 dimensions, including argument flow, evidence, audience framing, commercial integration, and voice, producing a 203-feature instrument with 176 structural features after removing style and format artifacts. An XGBoost classifier using structure alone achieved 97.0 macro-F1 on company-disjoint held-out data, compared with 88.1 for style-only features and 98.0 using all features. The structural signal survived self-rewriting: after each model reworded its own posts, performance fell only to 96.1 macro-F1, despite 73% of 13-word sequences being replaced. The recurring AI pattern was a tidy, self-announcing post that states its thesis early, previews its organization, contrasts modern practice with a legacy approach, uses confident institutional voice, and ends with a summary. Structure also supported source attribution: the classifier identified the correct human or AI model for 68.6% of posts, versus a 16.7% six-way chance rate. Human annotations validated the LLM scoring pipeline with human-human kappa of 0.939 and human-model kappa of 0.951, while the authors caution that the results cover single-pass generation and self-rewording, not all humanizer tools or collaborative editing.
Original abstract
Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice. We replicate StoryScope (Russell et al., 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models. A 214-feature instrument, applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.928, human-model 0.946), detects AI posts from its 187 structural features alone at 98.0 macro-F1 on held-out companies, unchanged (98.1) when every AI post is reworded by its own model. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 79.3% are attributed to the correct source against a 16.7% chance rate, and human posts occupy rare structural configurations. All effects replicate StoryScope's, consistent in direction and larger in magnitude. We release pipeline, instrument, prompts, code, and aggregate artifacts.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
Ej Zhou, Suchir Salhan, Catherine Arnett, Anna Korhonen
Independently trained language models may spontaneously learn representations that can be rotated and transferred across languages, suggesting multilingual capabilities without joint training.