Smaller Models, Better Rejects: Preference Distillation Scaling
AuthorsRui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao
Affiliations[
Resources
Smaller models can generate more useful negative examples than students themselves, making preference training for large language models cheaper and potentially better.
Key results
The study evaluates Qwen2.5 students from 7B to 72B.
For a 14B student, code avg@4 rises from 59.73% to 63.28% as Llama-3.2-1B-Instruct reject share reaches 100%.
Qwen2.5-3B-Instruct rejects improve the 14B student’s code avg@4 by 3.1 points.
Qwen2.5-3B-Instruct rejects improve the 14B student’s mathematical reasoning by 1.9 points.
Smaller-Base reject generation reduces the compute proxy by as much as 108.3-fold, with reductions ranging from 2.0-fold to 108.3-fold.
For the 1.5B source, lower-reference-likelihood selection improves avg@4 from 61.88% to 63.14%, a 1.26-point gain.
What the paper found
This paper revisits preference distillation’s default of using a student’s own failed responses as DPO rejects. In a SODA-style pipeline, Qwen2.5 students from 7B to 72B first receive sequence-level knowledge distillation, using GPT-4o-verified responses, then preference training on code and mathematics. Across KoDCode and OpenR1-Math-220k, rejects from smaller frozen Base models consistently outperform both vanilla Self and post-SeqKD Self rejects, even though they require less inference compute. For a 14B student, Qwen2.5-3B-Instruct rejects improve code avg@4 over vanilla Self by 3.1 points and mathematical reasoning by 1.9 points; replacing vanilla Self with Llama-3.2-1B-Instruct rejects raises code avg@4 from 59.73% to 63.28%. The cost advantage ranges from 2.0-fold to 108.3-fold. A linearized-feature analysis of DPO derives a finite-horizon utility bound and motivates three interventions: mix smaller-Base and Self rejects, reassign rejects across prompts or permute their code tokens, and select candidates with lower likelihood under the SeqKD reference. Real code structure remains useful even after prompt reassignment and token permutation, while likelihood-based reselection improves the 1.5B source from 61.88% to 63.14%, a 1.26-point gain. Math experiments use verifier-correct DeepSeek-R1 trajectories, and the resulting design principle is that effective rejects preserve task structure while remaining weakly coupled to the reference policy.
Original abstract
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.