NTH

Selecting Diverse SFT Traces Improves Post-RL Generalization

AuthorsDylan Zhang, Mingyuan Wu, Jinning Li

AffiliationsUniversity of Illinois Urbana-Champaign, work done at Google · Google

October 3, 2026 2 min read
Watch on YouTube
The one-line take

Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.

Key results

16.9%
OLMo3-7B RLVE gain

Post-RL pass@8 coverage improvement on environments held out from SFT.

6.2%
Single-model mathematics gain

Maximum mean pass@8 improvement across 10 mathematics benchmarks using Qwen3-4B-Thinking-2507 candidates.

54.7%
Diverse mixed-reward prompts

Share of mathematics prompts with mixed outcomes before RL.

46.9%
Similar mixed-reward prompts

Corresponding share for route-similar SFT.

3 hours
CPU selection runtime

Runtime for selecting from about 2.1M solutions on one CPU node.

What the paper found

This paper shows that verified reasoning traces are not interchangeable in the supervised fine-tuning, or SFT, stage before reinforcement learning with verifiable rewards. Instead of selecting similar solutions, it builds a rule-based topology fingerprint from reasoning-step types, transitions, branching, checking, and path statistics, then uses clustered farthest-point sampling to maximize route diversity at the same data budget. On RLVE, diverse SFT improved OLMo3-7B’s post-RL pass@8 coverage on environments held out from SFT by 16.9%, including problems harder than both training stages. The advantage was not dependent on multiple teachers: with Qwen3-4B-Thinking-2507 generating every candidate, diverse selection improved mean pass@8 across 10 mathematics benchmarks by up to 6.2%. The proposed mechanism is that route diversity places successful attempts within sampling reach on more prompts, producing mixed rewards that group-relative RL can learn from; before RL, mixed outcomes occurred on 54.7% of mathematics prompts versus 46.9% for similar SFT. The selector also worked on OpenThoughts3, INTELLECT-3, and Nemotron-Cascade 2, outperforming random, topology, gradient, embedding, and lexical baselines without model calls. On pools of about 2.1M solutions, it ran in about 3 hours on one CPU node, compared with 64–232 GPU-hours for more expensive alternatives. The results complement reasoning-model efforts from Qwen and DeepSeek by identifying route diversity, rather than SFT accuracy alone, as a practical way to prepare models for broader post-RL generalization.

Original abstract

Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →