NTH

Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families

AuthorsYifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, Jia Li

July 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper lets LLM distillation work across different tokenizers, making it easier and cheaper to transfer knowledge between otherwise incompatible model families.

Key results

35.1
Qwen3→Llama3.1 average

SFT baseline average on AIME24, AIME25, AIME26, MATH-500, GPQA-Diamond, and LiveCodeBench

40.7
Qwen3→Llama3.1 OPD average

Cross-tokenizer OPD average on the same benchmark suite

48.8
DeepSeek average

SFT baseline average for DeepSeek-R1→DeepSeek-R1-Distill-Qwen-7B

52.0
DeepSeek OPD average

Cross-tokenizer OPD average for DeepSeek-R1→DeepSeek-R1-Distill-Qwen-7B

20K
OPD data

Additional training prompts needed to reach 44.4% AIME24

4.8×10^19
OPD FLOPs

Total compute for OPD to reach 44.4% AIME24

What the paper found

Breaking the Tokenizer Barrier On Policy Distillation across Model Families, a collaboration involving The Hong Kong University of Science and Technology and Tencent, shows how to extend on-policy distillation beyond shared vocabularies. The paper argues that standard OPD, used in systems such as Qwen3, MiMo, GLM-5, and validated by Thinking Machines Lab, has been constrained by identical tokenizers, forcing weaker SFT-style cross-tokenizer transfer. Their solution is a dual-pointer chunk alignment algorithm, DPCA, that finds minimal synchronized text chunks across student and teacher token streams, plus a semantic-prior credit assignment rule that projects a teacher chunk’s log-likelihood back onto student tokens in closed form. In experiments with Qwen3-8B as teacher and Llama-3.1-8B-SFT as student, OPD lifts the average benchmark score from 35.1 to 40.7 across AIME24, AIME25, AIME26, MATH-500, GPQA-Diamond, and LiveCodeBench, outperforming ALM and CDM. In a second setting, DeepSeek-R1 to DeepSeek-R1-Distill-Qwen-7B, it raises the average from 48.8 to 52.0. The compute analysis is especially striking: to reach 44.4% on AIME24, OPD needs only 20K additional training samples and 4.8×10^19 FLOPs, versus 482K samples and 1.17×10^21 FLOPs for an SFT extrapolation, making OPD about 24.1× more compute-efficient. The key result is that tokenizer mismatch no longer blocks dense token-level teacher supervision.

Original abstract

On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models. However, existing OPD distillation methods require teacher and student models to share the same tokenizer, restricting the applicability of OPD within the model series. Current mainstream practice typically employs Supervised Fine-Tuning (SFT) on teacher-generated responses for cross-tokenizer distillation, which fails to capture the rich knowledge embedded in the teacher's probability distribution. In this work, we enable the standard on-policy distillation method to operate across model families, ensuring that high-fidelity token-level signals can propagate across different tokenizers with a precise token-mapping algorithm. Extensive experiments show that cross-tokenizer OPD is significantly more compute-efficient than baselines on various benchmarks. Our results unlock a broader range of teacher-student pairs for OPD, opening up new avenues for adapting and enhancing interactions between LLMs.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis