SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
AuthorsJinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao
SEED helps language-model agents learn from their own completed experiences by turning hindsight-generated skills into dense guidance during reinforcement learning.
Key results
SEED result with Qwen2.5-3B-Instruct on the seen ALFWorld benchmark.
SEED accuracy with Qwen2.5-3B-Instruct across seven search QA datasets.
SEED success rate with Qwen2.5-3B-Instruct.
SEED score using 60% of the training data, compared with GRPO’s 75.0% using the full dataset.
SEED macro-average success rate on unseen tasks, versus GRPO’s 70.9%.
What the paper found
SEED, from a Tsinghua University-led team, addresses sparse, trajectory-level rewards in long-horizon agentic reinforcement learning by converting completed on-policy episodes into natural-language hindsight skills. In Stage 1, the policy learns trajectory analysis from 1,440 annotated rollouts, with GLM-5.2 from Z.ai providing initial skill annotations. In Stage 2, the latest policy serves simultaneously as rollout actor and analyzer: it extracts reusable workflows or failure-avoidance rules, then re-scores the same sampled action tokens under ordinary and skill-augmented contexts. The detached log-probability shift becomes a confidence-gated, dense token-level on-policy distillation signal, jointly optimized with GRPO; skills are internalized during training and are absent at inference. Using Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, and Qwen3-1.7B-Instruct, SEED reaches 91.8 percent on ALFWorld, 45.7 percent on Search-based QA, and 78.9 percent exact success on WebShop. It also reaches 80.7 percent on ALFWorld using 60 percent of the training data, exceeding GRPO’s 75.0 percent with the full dataset, and obtains 86.2 percent on the ALFWorld unseen split versus GRPO’s 70.9 percent. On visual tasks with Qwen2.5-VL-3B-Instruct, SEED averages 91.0 percent across Sokoban and EZPoints, showing that synchronized hindsight distillation transfers beyond text-only agents.
Original abstract
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving framework that converts completed on-policy trajectories into training-time hindsight skills and distills their behavioral effect back into the policy model. SEED first fine-tunes the policy to analyze completed trajectories and generate natural-language skills that capture reusable workflows, decisive observations, or failure-avoidance rules. During RL, the current policy both collects trajectories and serves as the analyzer that extracts hindsight skills from them. Policy updates therefore improve subsequent decision making and skill analysis together, allowing hindsight supervision to evolve with the policy. SEED then re-scores the sampled actions under ordinary and skill-augmented contexts, converting the skill-induced probability shift into a dense token-level on-policy distillation signal. This signal is jointly optimized with outcome-based RL, keeping the auxiliary supervision aligned with the current trajectory distribution. Extensive experiments on text-based and vision-based agentic tasks show that SEED consistently improves performance and sample efficiency, exhibiting robust generalization to unseen scenarios. Our code is available at https://github.com/jinyangwu/SEED.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.