CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
AuthorsSatyam Kumar, Saurabh Jha
Resources
CADENCE helps small language models learn stronger mathematical reasoning from larger teachers using adaptive distillation, richer rewards, and modest hardware.
Key results
Performance of the 0.5B student after CADENCE distillation from Qwen2.5-Math-1.5B-Instruct.
Fraction of the Qwen2.5-Math-1.5B-Instruct to pretrained-student GSM8K gap closed.
Performance of the 0.5B student after distillation from Qwen2.5-3B-Instruct.
Fraction of the teacher-to-student GSM8K gap closed by CADENCE.
GSM8K pass@1 improvement in points for the 1.5B-to-0.5B configuration.
Approximate fraction of trajectories receiving nonzero reward after CCD numerical-proximity partial credit.
What the paper found
CADENCE, or Coverage-Adaptive On-Policy Distillation, targets three weaknesses in reasoning transfer: cold-start collapse, fixed divergence schedules, and sparse pass/fail rewards. Its DRIFT core trains on student-generated trajectories using per-token surrogate signals that mix forward and reverse KL, while COVA adapts the mixture according to the student’s measured coverage. FTB emphasizes high-entropy forking tokens, CCD grants numerical-proximity partial credit to incorrect-but-close math answers, LAP reinforces shorter correct solutions, EMR matches teacher and student entropy for calibration, and BSD performs correctness-gated self-distillation. Using Qwen2.5-Math-1.5B-Instruct or Qwen2.5-3B-Instruct teachers, CADENCE distills into a 0.5B Qwen2.5 student and evaluates on GSM8K and MATH-500 under a corrected 512-token protocol. With the 1.5B teacher, GSM8K pass@1 rises from 48.7% pretrained to 69.8%, closing 63.2% of the teacher gap; the 3B teacher reaches 72.1% and closes 76.2%. CADENCE exceeds the matched-compute DRIFT-plus-binary-reward baseline by 4.4 points on GSM8K, while increasing the nonzero-reward fraction to approximately 55%. The experiments use 5 seeds and run on a single Apple Mac Studio, showing that strong reasoning distillation can be achieved without datacenter-scale hardware. The authors emphasize that DRIFT optimizes practical per-token surrogates, not exact sequence-level KL gradients.
Original abstract
On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling, where time-only forward/reverse-KL interpolation ignores the student's coverage state; and (iii) binary reward sparsity, where pass/fail signals discard information from partially correct traces. We present CADENCE, a unified framework with a targeted fix for each. Its DRIFT mechanism schedules a per-token convex mixture of forward-KL and reverse-KL surrogate objectives on student-sampled trajectories (per-token surrogates, not sequence-level KL gradient estimators). Six components extend it: (A) COVA, a coverage-adaptive $β$ schedule accelerating the forward-to-reverse transition; (B) FTB, a forking-token boost concentrating gradient at high-entropy positions via a globally-normalized entropy reference; (C) CCD, a dense reward adding numerical-proximity partial credit for incorrect-but-close traces; (D) LAP, brevity-preferential correct-rollout reinforcement; (E) EMR, an entropy-matching calibration regularizer; (F) BSD, a bootstrapped self-distillation phase. On GSM8K and MATH-500 (corrected 512-token protocol, 5 seeds, reported std), CADENCE distills a 0.5B student from a 1.5B teacher to 69.8 $\pm$ 0.5% GSM8K pass@1 (from 48.7% pretrained; 63.2% of the teacher gap closed) and to 72.1 $\pm$ 0.4% with a 3B teacher (76.2% closed), beating the strongest matched-compute label-using baseline (DRIFT+binary reward) by +4.4 $\pm$ 0.7 points. All experiments run on a single Apple Mac Studio (M-series, 64GB unified memory), showing principled distillation reaches strong reasoning quality without datacenter-scale hardware.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.