Model Collapse as Cultural Evolution
AuthorsDongxin Guo, Jikun Wu, Siu Ming Yiu
Resources
This paper argues that model collapse in LLMs is really a kind of cultural evolution problem, and shows that self-training can first improve then damage compositional language structure before eventually degrading it.
Key results
The paper self-trains LLaMA-2-7B and Mistral-7B over 10 generations to study iterated learning effects.
Each generation uses 50,000 text passages as the bottlenecked training dataset.
The non-monotonic compositionality trajectory rises then falls under unfiltered self-training, with strong effect size support.
Bayesian evidence is reported for the non-monotonic compositionality trajectory in the main discriminative test.
LLM frequency-dependent regularization closely matches the human curve from Morgan and Levy (2016).
Distributional narrowing is observed over generations in LLaMA-2, with the Zipf exponent decreasing monotonically.
What the paper found
This paper argues that model collapse in LLM self-training is not just a statistical failure mode but an instance of cultural evolution under iterated learning, and it tests that claim by repeatedly fine-tuning LLaMA-2-7B and Mistral-7B for 10 generations on 50,000-passage datasets in English, German, and Turkish. The key novelty is a discriminative prediction: compositionality should first rise and then fall if compression operates without communicative grounding. That prediction is confirmed with strong effect sizes, including Hedges’ g = 1.87 and BF10 = 247 for the non-monotonic compositionality trajectory, and it persists even from a maximally regular PCFG seed, ruling out a noise-removal explanation. Task-grounded quality filtering, using an external evaluator on extractive QA, NLI, and summarization, is the only intervention that sustains compositionality; random 70% filtering behaves like no filtering. The paper also shows frequency-dependent regularization matching the human curve from Morgan and Levy (2016) with R2 = 0.94, monotonic morphological regularization across English, German, and Turkish, and distributional narrowing with Zipf exponent α falling from 1.07 to 0.82 in LLaMA-2 and 1.07 to 0.86 in Mistral. A dimension-specific loss order emerges as well: pragmatic constructions disappear before morphological ones, which disappear before core syntactic constructions, with entrenchment frequency correlating with survival generation at r = 0.73. Overall, the study reframes model collapse as a compression–communication tradeoff and suggests that self-training pipelines need task-grounded verification, not random subsampling, to preserve linguistic structure.
Original abstract
Model collapse, the progressive degradation of LLMs trained on their own outputs, has been characterized statistically but lacks a linguistic explanation for which structures degrade, in what order, and why. We show that iterated learning theory from cultural evolution fills this gap. We derive five falsifiable predictions, distinguish those uniquely discriminative for the theory from confirmatory ones, and test them by self-training LLaMA-2-7B and Mistral-7B over 10 generations in English, German, and Turkish. The critical discriminative finding: compositionality follows a non-monotonic trajectory (initially rising, then falling) under unfiltered self-training. This signature persists with maximally regular seed data (ruling out noise removal) and is sustained only by task-grounded filtering, not random filtering, providing the first LLM-scale evidence for the compression-communication tradeoff. All predictions are confirmed with large effect sizes (Hedges' $g > 1.6$; $\mathrm{BF}_{10} > 100$), and LLM regularization gradients closely match human behavioral data ($R^2 = 0.94$). These results reframe model collapse as a cultural transmission phenomenon and yield concrete principles for self-training pipeline design.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.