Reinforcement Learning from Rich Feedback with Distributional DAgger
AuthorsRishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad
Resources
This paper proposes DistIL, a new reinforcement-learning approach that turns rich feedback like traces and corrections into better policies with stronger theory and better performance than common self-distillation baselines.
Key results
DistIL final 5h average on scientific reasoning, compared with SDPO at 73.0.
DistIL final 5h average on scientific reasoning, compared with SDPO at 68.8.
DistIL coding performance with execution feedback.
DistIL exact-code correctness on coding evaluation.
Training problems used for hard mathematical reasoning experiments.
DistIL result versus OPSD 51.5 and SDPO 49.6.
What the paper found
Reinforcement Learning from Rich Feedback with Distributional DAgger, or DistIL, argues that the standard RLVR recipe in reasoning models wastes execution traces, critiques, and ground-truth solutions by reducing supervision to a terminal correctness bit. The paper, from the University of Southern California, proves two failure modes for prior on-policy self-distillation methods such as SDPO and OPSD: f-divergence objectives like reverse-KL and Jensen–Shannon do not guarantee monotonic policy improvement, and tokenwise local gradients can miss delayed credit assignment, with a two-step example where local updates converge to reward 1/3 while full sequence-level gradients reach 2/5. DistIL replaces that recipe with a distributional DAgger-style forward cross-entropy objective that can query black-box teachers and propagates future teacher-student disagreement back to earlier tokens. Under a local realizability condition, a natural-gradient step improves expected reward by ηΔ + O(η^2), and the online variant achieves a regret bound scaling as O(n^-1/4) for stochastic teachers or O(n^-1/2) when teacher variance is low. Empirically, on SciKnowEval L3, DistIL raises Qwen3-8B average Best@16 from 73.0 to 76.1 and Olmo3-7B-Instruct from 68.8 to 71.9 at 5 hours, on LiveCodeBench v6 it reaches Accuracy/Avg@16 0.656 and Score/Avg@16 0.482 versus SDPO’s 0.643 and 0.467, and on 738 hard math problems it improves Qwen3-4B AIME25 Avg@16 from 51.5 to 55.3 and Qwen3-8B from 69.7 to 71.1.
Original abstract
Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sample many responses and reward each with a single bit indicating whether the final answer is correct. Yet many settings provide rich feedback, including execution traces, tool outputs, expert corrections, and model self-evaluations. We study how to use such feedback through a distributional variant of the classic imitation learning algorithm DAgger, where the learner has local access to an expert distribution on states visited by the current policy. This yields a simple forward cross-entropy objective that admits a blackbox expert and whose sequence-level gradient {conduct rich credit assignment by propagating} future expert-student disagreement back to earlier decisions. We show that prior RL with self-distillation objectives based on reverse KL or Jensen-Shannon fail to guarantee monotonic policy improvement: even when the expert has higher reward, their updates may increase probability on worse actions. In contrast, we show that forward cross-entropy admits monotonic policy improvement and enjoys guarantees on regret. We further show that our objective optimizes a lower bound on teacher-weighted likelihood of success, leading to improved Pass@N. Empirically, our approach, DistIL, improves over RLVR and RL with self-distillation baselines across a variety of domains: scientific reasoning, coding, and solving hard mathematical problems.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.