NTH

Finetuning with Sampling: SFT Learns Better Than You Think

AuthorsAayush Karan, Sitan Chen, Yilun Du

AffiliationsHarvard University Website Code

October 9, 2026 2 min read
Watch on YouTube
The one-line take

By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.

Key results

66.0%
Chemistry accuracy, Sampling SFT

Qwen2.5-7B-Instruct on the SciKnowEval chemistry test set.

61.8%
Chemistry accuracy, OPSD

On-policy self-distillation baseline on the same Qwen2.5-7B-Instruct test.

49.5%
Hard MATH accuracy

Sampling SFT result for Qwen2.5-3B on MATH levels 3, 4, and 5.

58.2%
MATH500 accuracy

Sampling SFT result for Qwen2.5-3B.

10
MCMC steps

Number of projection-sampling steps used in the reported experiments.

What the paper found

This paper introduces projection sampling, a way to make supervised finetuning more effective without changing its training objective. Starting from expert demonstrations, a Metropolis–Hastings procedure repeatedly rewrites parts of each trace while preserving its solution, favoring candidates that are more likely under the base model. The resulting data moves closer to the model’s own distribution while retaining expert information, helping SFT learn new skills with less forgetting. On Qwen2.5-7B-Instruct, sampling SFT reached 66.0% chemistry accuracy, compared with 61.8% for on-policy self-distillation. On Qwen2.5-3B, it reached 49.5% on hard MATH problems and 58.2% on MATH500, outperforming the tested on-policy baselines on those measures. The approach also applies to Olmo-3-7B-Instruct; expert chemistry traces were generated with GPT-5. In the reported experiments, projection sampling used 10 MCMC steps, and the authors find that increasing sampling steps generally brings the data closer to the base model while improving downstream accuracy. The central implication is that data sampling can make ordinary SFT competitive with reinforcement learning and self-distillation, while preserving useful expert supervision.

Original abstract

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

Read the original paper

More in Large Language Models

Browse all 81 papers →
01Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
02Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis