On the Geometry of On-Policy Distillation
AuthorsZhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong, Zhilin Wang, Yanxu Chen, Rongduo Han, Sunbowen Lee, Yi R. Fung
Resources
This paper shows that on-policy distillation for language models follows its own unique path in parameter space, with updates quickly locking into a low-dimensional subspace that still preserves performance.
Key results
bf16-visible update sparsity in the controlled Qwen3-8B comparison
bf16-visible update sparsity in the controlled Qwen3-8B comparison
bf16-visible update sparsity in the controlled Qwen3-8B comparison
representative-layer principal-angle rotation is above this level
representative-layer principal-angle rotation is around this level
representative-layer principal-angle rotation is below this level
What the paper found
On the Geometry of On-Policy Distillation studies post-training dynamics in Qwen3 reasoning models and shows that on-policy distillation, or OPD, is neither supervised fine-tuning nor reinforcement learning with verifiable rewards, but a distinct geometry in parameter space. Using bf16-aware sparsity, principal-angle rotation, normalized spectral shift, and update-mask localization, the paper finds that OPD sits in a relaxed off-principal regime: in the controlled Qwen3-8B comparison, visible-update sparsity is 8.1% for SFT, 51.6% for OPD, and 77.2% for RLVR, while representative principal angles are above 10° for SFT, around 1° for OPD, and below 0.5° for RLVR. The authors then show subspace locking: OPD’s cumulative updates rapidly collapse into a narrow low-dimensional channel, and a rank-16 projection applied from about 20% of training preserves OPD performance but degrades SFT, indicating that the early subspace is functionally sufficient. Control experiments reveal that token sparsification at 25% or 50% and off-policy rollouts preserve this rank trajectory, but mixing OPD with RLVR changes it, so the critical lever is objective composition rather than runtime sampling. The study is framed around Qwen3-8B with a Qwen3-32B teacher on dapo-math-17k, and it argues for geometry-aware OPD designs that regulate update subspaces instead of only increasing token-level supervision density.
Original abstract
On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of OPD updates in parameter space and compare it with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). A suite of parameter-space diagnostics consistently places OPD in a relaxed off-principal regime: compared with SFT, its updates affect fewer weights and avoid principal directions more strongly, while compared with RLVR, they remain less tightly constrained. Beyond this static localization, OPD exhibits subspace locking: its cumulative updates rapidly enter a narrow low-dimensional channel. Constraining training to the update subspace formed early in training preserves OPD performance but substantially degrades SFT, indicating that the locked subspace is functionally sufficient for OPD. Control experiments further show that sparsifying the update tokens and shifting rollout generation off-policy preserve the rank dynamics, whereas mixing the OPD objective with RLVR changes them. Overall, these results suggest that OPD is not merely an intermediate point between SFT and RLVR, but induces its own update geometry in parameter space.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.