Finetuning with Sampling: SFT Learns Better Than You Think
AuthorsAayush Karan, Sitan Chen, Yilun Du
AffiliationsHarvard University Website Code
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Key results
Qwen2.5-7B-Instruct on the SciKnowEval chemistry test set.
On-policy self-distillation baseline on the same Qwen2.5-7B-Instruct test.
Sampling SFT result for Qwen2.5-3B on MATH levels 3, 4, and 5.
Sampling SFT result for Qwen2.5-3B.
Number of projection-sampling steps used in the reported experiments.
What the paper found
This paper introduces projection sampling, a way to make supervised finetuning more effective without changing its training objective. Starting from expert demonstrations, a Metropolis–Hastings procedure repeatedly rewrites parts of each trace while preserving its solution, favoring candidates that are more likely under the base model. The resulting data moves closer to the model’s own distribution while retaining expert information, helping SFT learn new skills with less forgetting. On Qwen2.5-7B-Instruct, sampling SFT reached 66.0% chemistry accuracy, compared with 61.8% for on-policy self-distillation. On Qwen2.5-3B, it reached 49.5% on hard MATH problems and 58.2% on MATH500, outperforming the tested on-policy baselines on those measures. The approach also applies to Olmo-3-7B-Instruct; expert chemistry traces were generated with GPT-5. In the reported experiments, projection sampling used 10 MCMC steps, and the authors find that increasing sampling steps generally brings the data closer to the base model while improving downstream accuracy. The central implication is that data sampling can make ordinary SFT competitive with reinforcement learning and self-distillation, while preserving useful expert supervision.
Original abstract
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
Read the original paperMore in Large Language Models
Browse all 81 papers →Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
LSPD uses reinforcement-learning ideas to make LLM policy distillation more sample-efficient while preserving the diversity needed for stronger reasoning.