NTH

CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

AuthorsSatyam Kumar, Saurabh Jha

August 5, 2026 2 min read
Watch on YouTube
The one-line take

CADENCE helps small language models learn stronger mathematical reasoning from larger teachers using adaptive distillation, richer rewards, and modest hardware.

Key results

69.8%
GSM8K pass@1 with 1.5B teacher

Performance of the 0.5B student after CADENCE distillation from Qwen2.5-Math-1.5B-Instruct.

63.2%
Teacher gap closed with 1.5B teacher

Fraction of the Qwen2.5-Math-1.5B-Instruct to pretrained-student GSM8K gap closed.

72.1%
GSM8K pass@1 with 3B teacher

Performance of the 0.5B student after distillation from Qwen2.5-3B-Instruct.

76.2%
Teacher gap closed with 3B teacher

Fraction of the teacher-to-student GSM8K gap closed by CADENCE.

4.4
Advantage over DRIFT plus binary reward

GSM8K pass@1 improvement in points for the 1.5B-to-0.5B configuration.

55%
Nonzero-reward fraction

Approximate fraction of trajectories receiving nonzero reward after CCD numerical-proximity partial credit.

What the paper found

CADENCE, or Coverage-Adaptive On-Policy Distillation, targets three weaknesses in reasoning transfer: cold-start collapse, fixed divergence schedules, and sparse pass/fail rewards. Its DRIFT core trains on student-generated trajectories using per-token surrogate signals that mix forward and reverse KL, while COVA adapts the mixture according to the student’s measured coverage. FTB emphasizes high-entropy forking tokens, CCD grants numerical-proximity partial credit to incorrect-but-close math answers, LAP reinforces shorter correct solutions, EMR matches teacher and student entropy for calibration, and BSD performs correctness-gated self-distillation. Using Qwen2.5-Math-1.5B-Instruct or Qwen2.5-3B-Instruct teachers, CADENCE distills into a 0.5B Qwen2.5 student and evaluates on GSM8K and MATH-500 under a corrected 512-token protocol. With the 1.5B teacher, GSM8K pass@1 rises from 48.7% pretrained to 69.8%, closing 63.2% of the teacher gap; the 3B teacher reaches 72.1% and closes 76.2%. CADENCE exceeds the matched-compute DRIFT-plus-binary-reward baseline by 4.4 points on GSM8K, while increasing the nonzero-reward fraction to approximately 55%. The experiments use 5 seeds and run on a single Apple Mac Studio, showing that strong reasoning distillation can be achieved without datacenter-scale hardware. The authors emphasize that DRIFT optimizes practical per-token surrogates, not exact sequence-level KL gradients.

Original abstract

On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling, where time-only forward/reverse-KL interpolation ignores the student's coverage state; and (iii) binary reward sparsity, where pass/fail signals discard information from partially correct traces. We present CADENCE, a unified framework with a targeted fix for each. Its DRIFT mechanism schedules a per-token convex mixture of forward-KL and reverse-KL surrogate objectives on student-sampled trajectories (per-token surrogates, not sequence-level KL gradient estimators). Six components extend it: (A) COVA, a coverage-adaptive $β$ schedule accelerating the forward-to-reverse transition; (B) FTB, a forking-token boost concentrating gradient at high-entropy positions via a globally-normalized entropy reference; (C) CCD, a dense reward adding numerical-proximity partial credit for incorrect-but-close traces; (D) LAP, brevity-preferential correct-rollout reinforcement; (E) EMR, an entropy-matching calibration regularizer; (F) BSD, a bootstrapped self-distillation phase. On GSM8K and MATH-500 (corrected 512-token protocol, 5 seeds, reported std), CADENCE distills a 0.5B student from a 1.5B teacher to 69.8 $\pm$ 0.5% GSM8K pass@1 (from 48.7% pretrained; 63.2% of the teacher gap closed) and to 72.1 $\pm$ 0.4% with a 3B teacher (76.2% closed), beating the strongest matched-compute label-using baseline (DRIFT+binary reward) by +4.4 $\pm$ 0.7 points. All experiments run on a single Apple Mac Studio (M-series, 64GB unified memory), showing principled distillation reaches strong reasoning quality without datacenter-scale hardware.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis