NTH

AsyncOPD: How Stale Can On-Policy Distillation Be?

AuthorsWonjun Kang, Kevin Galim, Seunghyuk Oh, Minjun Kang, Sanghyun Park, Donghoon Kim, Minjae Lee, Minseo Kim, Rishabh Tiwari, Yuchen Zeng, Hyung Il Koo, Kangwook Lee

July 1, 2026 2 min read
Watch on YouTube
The one-line take

AsyncOPD shows how to train language models faster with asynchronous on-policy distillation without losing accuracy, by carefully handling stale rollouts and teacher-feedback estimates.

Key results

3.8x
training throughput gain

AsyncOPD speedup over strict synchronous training on Qwen3-Base/Qwen3 students

1.6x
training throughput gain

Lower end of AsyncOPD speedup over strict synchronous training

0.0149
MC64 variance ratio

Fixed-prefix Monte Carlo variance at 64 samples relative to one-sample baseline

What the paper found

AsyncOPD: How Stale Can On-Policy Distillation Be? is the first systematic study of staleness in asynchronous on-policy distillation for large language model post-training, using a Qwen3-30B-A3B-Instruct-2507 teacher and Qwen3-4B-Base, 1.7B, and 8B students on DeepMath. The paper shows that KL direction fundamentally changes the stale-data problem: forward KL is teacher-weighted and stays comparatively robust as rollout lag grows, while reverse KL is student-weighted and degrades faster. For reverse KL, the strongest correction is not a complex asynchronous-RL surrogate such as Decoupled PPO or M2PO, but a simpler OPD-specific estimator that recomputes the advantage Aθ at learner time and removes PPO clipping. The authors also show that stale top-k teacher caches are support-mismatched, whereas sampled Monte Carlo reverse-KL remains correctable by old-to-current importance sampling; increasing the number of local samples from 1 to 64 sharply reduces variance, with fixed-prefix variance falling to 0.0149 of the one-sample baseline. Building on these estimator choices, AsyncOPD overlaps rollout, teacher scoring, and learning in a fully asynchronous pipeline and improves throughput by 1.6× to 3.8× over strict synchronous training while maintaining comparable final accuracy on AIME24, AIME25, and AMC.

Original abstract

On-policy distillation (OPD) trains a student on its own rollouts guided by teacher feedback and is becoming increasingly important for large language model (LLM) post-training. Like reinforcement learning (RL), however, OPD faces an on-policy systems bottleneck, as rollouts can dominate training time for reasoning workloads. Asynchronous training pipelines can alleviate this bottleneck by decoupling rollout generation from learner updates, but doing so introduces stale-policy data. While prior work has studied stale data in asynchronous RL, its effects in OPD remain underexplored. We present the first systematic study of staleness in asynchronous OPD, focusing on a practical setting where teacher feedback is implemented through local KL losses and full-vocabulary teacher logits are too expensive to store or transfer, necessitating finite teacher-score caches. We first show that KL direction changes the stale-data problem: teacher-weighted forward KL is more robust to stale rollouts, whereas student-weighted reverse KL is vulnerable. Second, for this vulnerable reverse-KL case, we study whether methods designed to stabilize asynchronous RL can mitigate OPD staleness. In our experiments, they do not improve over a simpler OPD-specific surrogate: recomputing the reverse-KL signal under the current student at learner time. Third, we analyze how finite teacher-score caches create a bias-variance tradeoff for sparse and sampled reverse-KL OPD estimators. This motivates multi-sample Monte Carlo (MC), which preserves MC correctability while reducing one-sample variance. Finally, we present and open-source AsyncOPD, a fully asynchronous OPD training pipeline built from these estimator choices. Experiments show that AsyncOPD improves training throughput by $1.6\times$ to $3.8\times$ over strict synchronous training while reaching comparable accuracy.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis