NTH

Reinforcement Learning from Rich Feedback with Distributional DAgger

AuthorsRishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

June 11, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes DistIL, a new reinforcement-learning approach that turns rich feedback like traces and corrections into better policies with stronger theory and better performance than common self-distillation baselines.

Key results

76.1
SciKnowEval L3 Qwen3-8B Average Best@16

DistIL final 5h average on scientific reasoning, compared with SDPO at 73.0.

71.9
SciKnowEval L3 Olmo3-7B-Instruct Average Best@16

DistIL final 5h average on scientific reasoning, compared with SDPO at 68.8.

0.656
LiveCodeBench v6 Accuracy/Avg@16

DistIL coding performance with execution feedback.

0.482
LiveCodeBench v6 Score/Avg@16

DistIL exact-code correctness on coding evaluation.

738
Hard math training set size

Training problems used for hard mathematical reasoning experiments.

55.3
AIME25 Qwen3-4B Avg@16

DistIL result versus OPSD 51.5 and SDPO 49.6.

What the paper found

Reinforcement Learning from Rich Feedback with Distributional DAgger, or DistIL, argues that the standard RLVR recipe in reasoning models wastes execution traces, critiques, and ground-truth solutions by reducing supervision to a terminal correctness bit. The paper, from the University of Southern California, proves two failure modes for prior on-policy self-distillation methods such as SDPO and OPSD: f-divergence objectives like reverse-KL and Jensen–Shannon do not guarantee monotonic policy improvement, and tokenwise local gradients can miss delayed credit assignment, with a two-step example where local updates converge to reward 1/3 while full sequence-level gradients reach 2/5. DistIL replaces that recipe with a distributional DAgger-style forward cross-entropy objective that can query black-box teachers and propagates future teacher-student disagreement back to earlier tokens. Under a local realizability condition, a natural-gradient step improves expected reward by ηΔ + O(η^2), and the online variant achieves a regret bound scaling as O(n^-1/4) for stochastic teachers or O(n^-1/2) when teacher variance is low. Empirically, on SciKnowEval L3, DistIL raises Qwen3-8B average Best@16 from 73.0 to 76.1 and Olmo3-7B-Instruct from 68.8 to 71.9 at 5 hours, on LiveCodeBench v6 it reaches Accuracy/Avg@16 0.656 and Score/Avg@16 0.482 versus SDPO’s 0.643 and 0.467, and on 738 hard math problems it improves Qwen3-4B AIME25 Avg@16 from 51.5 to 55.3 and Qwen3-8B from 69.7 to 71.1.

Original abstract

Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sample many responses and reward each with a single bit indicating whether the final answer is correct. Yet many settings provide rich feedback, including execution traces, tool outputs, expert corrections, and model self-evaluations. We study how to use such feedback through a distributional variant of the classic imitation learning algorithm DAgger, where the learner has local access to an expert distribution on states visited by the current policy. This yields a simple forward cross-entropy objective that admits a blackbox expert and whose sequence-level gradient {conduct rich credit assignment by propagating} future expert-student disagreement back to earlier decisions. We show that prior RL with self-distillation objectives based on reverse KL or Jensen-Shannon fail to guarantee monotonic policy improvement: even when the expert has higher reward, their updates may increase probability on worse actions. In contrast, we show that forward cross-entropy admits monotonic policy improvement and enjoys guarantees on regret. We further show that our objective optimizes a lower bound on teacher-weighted likelihood of success, leading to improved Pass@N. Empirically, our approach, DistIL, improves over RLVR and RL with self-distillation baselines across a variety of domains: scientific reasoning, coding, and solving hard mathematical problems.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →