ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
AuthorsQingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao
Resources
ShortOPD helps pruned language models regain useful free-form generation by training first on the prefixes they can handle and gradually extending their rollout length.
Key results
Fraction of Qwen3-4B-Instruct parameters removed before recovery.
Prompts spanning math, code, and open-ended instruction domains.
Eight-task normalized generation score after one epoch.
Eight-task normalized score of the 25%-pruned model without recovery.
End-to-end wall-clock time for dynamic-budget recovery.
Wall-clock time for the fixed long-rollout comparison.
What the paper found
ShortOPD, from researchers at ByteDance and the Chinese Academy of Sciences, addresses a failure mode in structured pruning: compressed models can retain correct trajectories in their sampling distribution while greedy free-form generation collapses into repetition. Using Block-Influence pruning on Qwen3-4B-Instruct, the method keeps the pre-pruning model frozen as a teacher and applies on-policy distillation with dense token-level generalized Jensen–Shannon targets, using the top-100 logits plus tail mass. Its novelty is a repetition-gated short-to-long rollout controller: it detects terminal periodic suffixes, treats the surviving prefix as the effective length, and uses exponential moving averages to shrink or expand future budgets. On a 45,447-prompt corpus spanning GSM8K, MATH, NVIDIA OpenCodeInstruct, MBPP, ShareGPT-style instructions, and UltraChat-200K, 25% pruning reduced the eight-task generation average to 5.71, while one epoch of ShortOPD restored it to 48.46, compared with 30.52 for conventional teacher-forced KD; evaluation included GPT-5.5 judging for open-ended tasks. ShortOPD reached this quality in 8.5 hours, versus 35.9 hours for a fixed 8192-token horizon, using 250M rollout tokens rather than 869M, while staying within two average-score points of the longer schedule. The study argues that post-pruning recovery should target the compressed model’s own states with dense teacher supervision and dynamically allocate compute as generation quality improves.
Original abstract
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OPD schedule that detects teacher-confirmed repetitive suffixes, treats the surviving prefix as each rollout's effective length, and allocates future rollout budgets to the effective lengths the policy can currently use. Across math, code, and open-ended generation, \shortopd\ raises the compressed model's score to about $9\times$ its unrecovered value and $1.6$--$4.4\times$ standard recovery recipes (SFT w/o KD, KD, and SeqKD), and it matches a fixed $8192$-token rollout horizon within two points using a quarter of the training time ($8.5$ vs.\ $35.9$ hours) and $71\%$ fewer rollout tokens. We hope this recipe helps move structured pruning beyond marginal gains on perplexity and multiple-choice benchmarks, a step closer to deployment-ready generation quality.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.