NTH

On the Geometry of On-Policy Distillation

AuthorsZhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong, Zhilin Wang, Yanxu Chen, Rongduo Han, Sunbowen Lee, Yi R. Fung

July 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that on-policy distillation for language models follows its own unique path in parameter space, with updates quickly locking into a low-dimensional subspace that still preserves performance.

Key results

8.1%
SFT sparsity

bf16-visible update sparsity in the controlled Qwen3-8B comparison

51.6%
OPD sparsity

bf16-visible update sparsity in the controlled Qwen3-8B comparison

77.2%
RLVR sparsity

bf16-visible update sparsity in the controlled Qwen3-8B comparison

10
SFT principal angle

representative-layer principal-angle rotation is above this level

1
OPD principal angle

representative-layer principal-angle rotation is around this level

0.5
RLVR principal angle

representative-layer principal-angle rotation is below this level

What the paper found

On the Geometry of On-Policy Distillation studies post-training dynamics in Qwen3 reasoning models and shows that on-policy distillation, or OPD, is neither supervised fine-tuning nor reinforcement learning with verifiable rewards, but a distinct geometry in parameter space. Using bf16-aware sparsity, principal-angle rotation, normalized spectral shift, and update-mask localization, the paper finds that OPD sits in a relaxed off-principal regime: in the controlled Qwen3-8B comparison, visible-update sparsity is 8.1% for SFT, 51.6% for OPD, and 77.2% for RLVR, while representative principal angles are above 10° for SFT, around 1° for OPD, and below 0.5° for RLVR. The authors then show subspace locking: OPD’s cumulative updates rapidly collapse into a narrow low-dimensional channel, and a rank-16 projection applied from about 20% of training preserves OPD performance but degrades SFT, indicating that the early subspace is functionally sufficient. Control experiments reveal that token sparsification at 25% or 50% and off-policy rollouts preserve this rank trajectory, but mixing OPD with RLVR changes it, so the critical lever is objective composition rather than runtime sampling. The study is framed around Qwen3-8B with a Qwen3-32B teacher on dapo-math-17k, and it argues for geometry-aware OPD designs that regulate update subspaces instead of only increasing token-level supervision density.

Original abstract

On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of OPD updates in parameter space and compare it with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). A suite of parameter-space diagnostics consistently places OPD in a relaxed off-principal regime: compared with SFT, its updates affect fewer weights and avoid principal directions more strongly, while compared with RLVR, they remain less tightly constrained. Beyond this static localization, OPD exhibits subspace locking: its cumulative updates rapidly enter a narrow low-dimensional channel. Constraining training to the update subspace formed early in training preserves OPD performance but substantially degrades SFT, indicating that the locked subspace is functionally sufficient for OPD. Control experiments further show that sparsifying the update tokens and shifting rollout generation off-policy preserve the rank dynamics, whereas mixing the OPD objective with RLVR changes them. Overall, these results suggest that OPD is not merely an intermediate point between SFT and RLVR, but induces its own update geometry in parameter space.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis