Strategically Diverse Sampling for Self-Training
AuthorsAlexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
AffiliationsSchool of Informatics, University of Edinburgh · {alex.gurung,esmeralda.whitammer}@ed.ac.uk, mlap@inf.ed.ac.uk
Resources
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
Key results
Qwen3-4B-Instruct on LiveCodeBench and OJBench frontier problems.
Strategic sampling outperformed IID-4 at 5.2% with the same four-sample budget.
Coverage at the 15% perplexity-improvement threshold, versus 1.97% for the base model and 2.08% for IID.
Held-out Macro-Avg score for IID distillation, below self-generated VS-4 at 22.8%.
Held-out Macro-Avg pass@64 after reinforcement learning, versus 16.0% for VS and 10.9% for IID-4.
What the paper found
This paper argues that self-training improves more from strategic diversity than from simply collecting more correct answers. Instead of IID sampling, it introduces G ROOT, which builds a hierarchy of solution strategies and samples distinct root-to-leaf paths, and an approach-level version of Verbalized Sampling, or VS, which generates an unstructured set of alternatives before producing responses. Experiments with Qwen3-4B-Instruct on Cobalt, LiveCodeBench, OJBench, and Next-Chapter Prediction show that four strategically sampled traces can outperform 64 IID traces: on held-out coding problems, G ROOT-4 and VS-4 reach Macro-Avg pass@64 scores of 15.5% and 17.5%, versus 5.2% for IID-4. On Next-Chapter Prediction, VS-trained models reach 4.08% coverage@32 at the 15% perplexity-improvement threshold, compared with 1.97% for the base model and 2.08% for IID training. The effect is driven by approach diversity rather than correctness alone: strategically diverse but incorrect traces still beat IID training. Self-generated strategic data also surpasses IID distillation from a 235B teacher, with VS-4 reaching 22.8% held-out Macro-Avg pass@64 versus 13.4% for teacher-generated IID data. Strategic initialization further benefits downstream optimization: after RL, G ROOT reaches 18.0% held-out Macro-Avg pass@64, compared with 16.0% for VS and 10.9% for IID-4. Replication with NVIDIA’s Nemotron3-Nano-4B supports the same trend, suggesting that preserving distinct reasoning strategies is a stronger lever than temperature increases, correctness filtering, or teacher scale.
Original abstract
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Hui Wei, Licai Sun, Guoying Zhao
Human-JEPA aims to help vision systems understand people now and predict what they will do next using one efficient self-supervised video model.