NTH

Strategically Diverse Sampling for Self-Training

AuthorsAlexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

AffiliationsSchool of Informatics, University of Edinburgh · {alex.gurung,esmeralda.whitammer}@ed.ac.uk, mlap@inf.ed.ac.uk

September 30, 2026 2 min read
Watch on YouTube
The one-line take

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Key results

15.5%
G ROOT-4 held-out Macro-Avg pass@64

Qwen3-4B-Instruct on LiveCodeBench and OJBench frontier problems.

17.5%
VS-4 held-out Macro-Avg pass@64

Strategic sampling outperformed IID-4 at 5.2% with the same four-sample budget.

4.08%
VS NCP coverage@32

Coverage at the 15% perplexity-improvement threshold, versus 1.97% for the base model and 2.08% for IID.

13.4%
235B teacher IID pass@64

Held-out Macro-Avg score for IID distillation, below self-generated VS-4 at 22.8%.

18.0%
G ROOT-4 after RL

Held-out Macro-Avg pass@64 after reinforcement learning, versus 16.0% for VS and 10.9% for IID-4.

What the paper found

This paper argues that self-training improves more from strategic diversity than from simply collecting more correct answers. Instead of IID sampling, it introduces G ROOT, which builds a hierarchy of solution strategies and samples distinct root-to-leaf paths, and an approach-level version of Verbalized Sampling, or VS, which generates an unstructured set of alternatives before producing responses. Experiments with Qwen3-4B-Instruct on Cobalt, LiveCodeBench, OJBench, and Next-Chapter Prediction show that four strategically sampled traces can outperform 64 IID traces: on held-out coding problems, G ROOT-4 and VS-4 reach Macro-Avg pass@64 scores of 15.5% and 17.5%, versus 5.2% for IID-4. On Next-Chapter Prediction, VS-trained models reach 4.08% coverage@32 at the 15% perplexity-improvement threshold, compared with 1.97% for the base model and 2.08% for IID training. The effect is driven by approach diversity rather than correctness alone: strategically diverse but incorrect traces still beat IID training. Self-generated strategic data also surpasses IID distillation from a 235B teacher, with VS-4 reaching 22.8% held-out Macro-Avg pass@64 versus 13.4% for teacher-generated IID data. Strategic initialization further benefits downstream optimization: after RL, G ROOT reaches 18.0% held-out Macro-Avg pass@64, compared with 16.0% for VS and 10.9% for IID-4. Replication with NVIDIA’s Nemotron3-Nano-4B supports the same trend, suggesting that preserving distinct reasoning strategies is a stronger lever than temperature increases, correctness filtering, or teacher scale.

Original abstract

Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis