How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
AuthorsJenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AffiliationsUniversity of Maryland · Pangram Labs
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
Key results
Share of web tokens labeled AI-generated after FineWeb filtering.
Number of pretrained models evaluated across sizes and AI-to-human token ratios.
Released corpus size in tokens with AI, topic, and format labels.
Paired RMSE ×10^-3 on held-out model sizes for the proposed scaling law.
Compute required at the August 2026 AI share relative to training on the human-only subset.
Harmful runs incorrectly reported as improvements by a mixed validation set containing 22.3% AI text.
What the paper found
This study examines “wild AI text”: unlabeled web content produced by many systems for human readers, rather than curated synthetic data or recursive model-collapse data. In the ChatGPT era, Pangram labels 31.1% of August 2026 web tokens as AI-generated, while quality pipelines such as FineWeb and DCLM preferentially retain it. The researchers trained 800 nanochat language models ranging from 19.9M to 973M parameters, varying AI-to-human token ratios, and evaluated loss on C4, FineWeb, Paloma, and Cosmopedia. AI tokens help when human data is scarce, but their value saturates and becomes negative near the Chinchilla budget of 20 human tokens per parameter; they remain highly valuable when the target is AI text. The proposed scaling law combines a saturating benefit term with a logarithmic harm penalty and recovers Chinchilla when AI data is absent. Fitted on models up to 268M parameters, it predicts held-out models up to 3.6× larger with a paired C4 error of 0.83 × 10^-3, outperforming existing laws. WildAI, an 83B-token corpus labeled with AI, topic, and format metadata, supports the analysis; EditLens Llama-3.2-3B and Pangram 3.3.2 provide detection. Practically, training on unfiltered August data requires 1.6× the compute of its human-only subset, so the paper recommends filtering AI text for human-language modeling, repeating human data before adding AI data, and reporting human and AI validation losses separately. Mixed validation hides harm in 95.5% of harmful runs.
Original abstract
Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.
Read the original paperMore in Foundation Models
Browse all 47 papers →TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.
nnFoundation: 3D Foundation Models for Radiology
Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, Bálint Kovács, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica Kächele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Krämer, Maximilian Zenk, Tim Rädsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David Kügler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina ...
nnFoundation trains complementary 3D convolutional and transformer models on 2.1 million medical scans and shows that matching architecture to the task can improve transfer across diverse radiology applications.