Self-Play Pretraining with Zero Data
AuthorsAditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
AffiliationsIndependent Researcher · Tel Aviv University · Stanford University · LAPTh, USMB
Resources
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Key results
Experiments used models below this parameter count.
Self-play discovered Fibonacci, geometric, quadratic, and cubic sequence families by this round.
REVERSE STRING, STACK, and ASSOCIATIVE RECALL approached this accuracy after sufficient demonstrations.
Parameter count of the model used in downstream natural-data pretraining.
Natural-data tokens required with self-play initialization, versus 496M from random initialization.
What the paper found
Self-Play Pretraining with Zero Data tests whether a language model can discover its own training distribution without human-curated or natural data. Starting from random initialization, two Meta Llama-style decoder-only transformers co-evolve: a generator writes programs in a Brainf*ck-like universal Turing-complete language, executes them to produce byte sequences, and a learner trains on those outputs with next-token cross-entropy. Reinforcement learning steers the generator using a preconditioned gradient-alignment reward, favoring programs that produce learnable but not-yet-mastered structure; replay, mutation, and KL regularization stabilize the curriculum. With models below 25M parameters and a 4K context, zero-shot loss scales predictably with self-play compute across DCLM text, CIFAR-10 images, Mutopia melodies, audio, speech, DNA, mathematics, and code, despite no evaluation data entering gradient updates. The generator discovers Fibonacci, geometric, quadratic, and cubic sequences by round 512, while a fixed universal-program prior required more than 53,000 rounds to encounter comparable families. Learners also develop in-context capabilities: after sufficient demonstrations, they reach almost 100% exact-match accuracy on REVERSE STRING, STACK, and ASSOCIATIVE RECALL, as well as mathematical MAX, MIN, and SUM tasks. As a practical pre-pretraining result, a 24.4M-parameter self-play checkpoint reduces ESC-50 convergence from 496M natural-data tokens to 320M, suggesting that self-play learns transferable predictive structure rather than factual knowledge; it complements, rather than replaces, natural-data pretraining.
Original abstract
Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Hui Wei, Licai Sun, Guoying Zhao
Human-JEPA aims to help vision systems understand people now and predict what they will do next using one efficient self-supervised video model.