NTH

Smaller Models, Better Rejects: Preference Distillation Scaling

AuthorsRui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao

Affiliations[

October 8, 2026 2 min read
Watch on YouTube
The one-line take

Smaller models can generate more useful negative examples than students themselves, making preference training for large language models cheaper and potentially better.

Key results

72B
Student scaling range

The study evaluates Qwen2.5 students from 7B to 72B.

63.28%
Llama reject mixture endpoint

For a 14B student, code avg@4 rises from 59.73% to 63.28% as Llama-3.2-1B-Instruct reject share reaches 100%.

3.1
Code improvement over vanilla Self

Qwen2.5-3B-Instruct rejects improve the 14B student’s code avg@4 by 3.1 points.

1.9
Math improvement over vanilla Self

Qwen2.5-3B-Instruct rejects improve the 14B student’s mathematical reasoning by 1.9 points.

108.3
Maximum inference-cost reduction

Smaller-Base reject generation reduces the compute proxy by as much as 108.3-fold, with reductions ranging from 2.0-fold to 108.3-fold.

1.26
Likelihood-reselection gain

For the 1.5B source, lower-reference-likelihood selection improves avg@4 from 61.88% to 63.14%, a 1.26-point gain.

What the paper found

This paper revisits preference distillation’s default of using a student’s own failed responses as DPO rejects. In a SODA-style pipeline, Qwen2.5 students from 7B to 72B first receive sequence-level knowledge distillation, using GPT-4o-verified responses, then preference training on code and mathematics. Across KoDCode and OpenR1-Math-220k, rejects from smaller frozen Base models consistently outperform both vanilla Self and post-SeqKD Self rejects, even though they require less inference compute. For a 14B student, Qwen2.5-3B-Instruct rejects improve code avg@4 over vanilla Self by 3.1 points and mathematical reasoning by 1.9 points; replacing vanilla Self with Llama-3.2-1B-Instruct rejects raises code avg@4 from 59.73% to 63.28%. The cost advantage ranges from 2.0-fold to 108.3-fold. A linearized-feature analysis of DPO derives a finite-horizon utility bound and motivates three interventions: mix smaller-Base and Self rejects, reassign rejects across prompts or permute their code tokens, and select candidates with lower likelihood under the SeqKD reference. Real code structure remains useful even after prompt reassignment and token permutation, while likelihood-based reselection improves the 1.5B source from 61.88% to 63.14%, a 1.26-point gain. Math experiments use verifier-correct DeepSeek-R1 trajectories, and the resulting design principle is that effective rejects preserve task structure while remaining weakly coupled to the reference policy.

Original abstract

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis