Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families
AuthorsYifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, Jia Li
Resources
This paper lets LLM distillation work across different tokenizers, making it easier and cheaper to transfer knowledge between otherwise incompatible model families.
Key results
SFT baseline average on AIME24, AIME25, AIME26, MATH-500, GPQA-Diamond, and LiveCodeBench
Cross-tokenizer OPD average on the same benchmark suite
SFT baseline average for DeepSeek-R1→DeepSeek-R1-Distill-Qwen-7B
Cross-tokenizer OPD average for DeepSeek-R1→DeepSeek-R1-Distill-Qwen-7B
Additional training prompts needed to reach 44.4% AIME24
Total compute for OPD to reach 44.4% AIME24
What the paper found
Breaking the Tokenizer Barrier On Policy Distillation across Model Families, a collaboration involving The Hong Kong University of Science and Technology and Tencent, shows how to extend on-policy distillation beyond shared vocabularies. The paper argues that standard OPD, used in systems such as Qwen3, MiMo, GLM-5, and validated by Thinking Machines Lab, has been constrained by identical tokenizers, forcing weaker SFT-style cross-tokenizer transfer. Their solution is a dual-pointer chunk alignment algorithm, DPCA, that finds minimal synchronized text chunks across student and teacher token streams, plus a semantic-prior credit assignment rule that projects a teacher chunk’s log-likelihood back onto student tokens in closed form. In experiments with Qwen3-8B as teacher and Llama-3.1-8B-SFT as student, OPD lifts the average benchmark score from 35.1 to 40.7 across AIME24, AIME25, AIME26, MATH-500, GPQA-Diamond, and LiveCodeBench, outperforming ALM and CDM. In a second setting, DeepSeek-R1 to DeepSeek-R1-Distill-Qwen-7B, it raises the average from 48.8 to 52.0. The compute analysis is especially striking: to reach 44.4% on AIME24, OPD needs only 20K additional training samples and 4.8×10^19 FLOPs, versus 482K samples and 1.17×10^21 FLOPs for an SFT extrapolation, making OPD about 24.1× more compute-efficient. The key result is that tokenizer mismatch no longer blocks dense token-level teacher supervision.
Original abstract
On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models. However, existing OPD distillation methods require teacher and student models to share the same tokenizer, restricting the applicability of OPD within the model series. Current mainstream practice typically employs Supervised Fine-Tuning (SFT) on teacher-generated responses for cross-tokenizer distillation, which fails to capture the rich knowledge embedded in the teacher's probability distribution. In this work, we enable the standard on-policy distillation method to operate across model families, ensuring that high-fidelity token-level signals can propagate across different tokenizers with a precise token-mapping algorithm. Extensive experiments show that cross-tokenizer OPD is significantly more compute-efficient than baselines on various benchmarks. Our results unlock a broader range of teacher-student pairs for OPD, opening up new avenues for adapting and enhancing interactions between LLMs.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.