Selecting Diverse SFT Traces Improves Post-RL Generalization
AuthorsDylan Zhang, Mingyuan Wu, Jinning Li
AffiliationsUniversity of Illinois Urbana-Champaign, work done at Google · Google
Resources
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Key results
Post-RL pass@8 coverage improvement on environments held out from SFT.
Maximum mean pass@8 improvement across 10 mathematics benchmarks using Qwen3-4B-Thinking-2507 candidates.
Share of mathematics prompts with mixed outcomes before RL.
Corresponding share for route-similar SFT.
Runtime for selecting from about 2.1M solutions on one CPU node.
What the paper found
This paper shows that verified reasoning traces are not interchangeable in the supervised fine-tuning, or SFT, stage before reinforcement learning with verifiable rewards. Instead of selecting similar solutions, it builds a rule-based topology fingerprint from reasoning-step types, transitions, branching, checking, and path statistics, then uses clustered farthest-point sampling to maximize route diversity at the same data budget. On RLVE, diverse SFT improved OLMo3-7B’s post-RL pass@8 coverage on environments held out from SFT by 16.9%, including problems harder than both training stages. The advantage was not dependent on multiple teachers: with Qwen3-4B-Thinking-2507 generating every candidate, diverse selection improved mean pass@8 across 10 mathematics benchmarks by up to 6.2%. The proposed mechanism is that route diversity places successful attempts within sampling reach on more prompts, producing mixed rewards that group-relative RL can learn from; before RL, mixed outcomes occurred on 54.7% of mathematics prompts versus 46.9% for similar SFT. The selector also worked on OpenThoughts3, INTELLECT-3, and Nemotron-Cascade 2, outperforming random, topology, gradient, embedding, and lexical baselines without model calls. On pools of about 2.1M solutions, it ran in about 3 hours on one CPU node, compared with 64–232 GPU-hours for more expensive alternatives. The results complement reasoning-model efforts from Qwen and DeepSeek by identifying route diversity, rather than SFT accuracy alone, as a practical way to prepare models for broader post-RL generalization.
Original abstract
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville
NGU helps RL-trained LLMs stop over-practicing easy problems and spend more effort solving the hard ones.