PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
AuthorsYang Tian, Rui Wang, Xumeng Wen, Junjie Li, Shizhao Sun, Lei Song, Jiang Bian, Bo Zhao
Resources
PBSD turns sparse final rewards into step-by-step learning signals for agents by using Bayesian self-distillation to figure out which actions really helped.
Key results
PBSD in-domain validation accuracy
PBSD performance on the stratified BrowseComp subset
PBSD improvement over GRPO on validation
PBSD improvement over GRPO on BC(300)
Best trained-agent BrowseComp score reported in the benchmark table
What the paper found
PBSD, or Privileged Bayesian Self-Distillation, is a long-horizon credit-assignment method for reinforcement learning with verifiable rewards that targets the core weakness of outcome-only training: a single final reward cannot tell which of hundreds of search and tool-use turns were actually useful. The paper, from Shanghai Jiao Tong University and XYZ AI Lab, reformulates trajectory quality as a posterior-to-prior evidence ratio for the verified answer, then uses Bayes’ rule to convert that intractable quantity into a turn-level likelihood ratio between a standard student policy and an answer-conditioned privileged teacher. This Bayesian evidence score is autoregressively decomposed into per-turn signals and used only as a detached weight on GRPO-style advantages, avoiding direct imitation and information leakage. The method is evaluated on Qwen3-30B-A3B-Thinking-2507 with 64K training context and up to 300 interaction turns, using 7.5K supervised trajectories and 575 RL training examples from a synthetic Wikipedia/OpenSeeker pipeline. On the in-domain validation set and a stratified 300-example BrowseComp subset, PBSD outperforms GRPO by 2.62 and 3.50 points, reaching 40.87 and 35.83 respectively, and it sets the best trained-agent BrowseComp score in the broader benchmark table at 46.21 with SFT+PBSD. Ablations show that replay-free evidence scoring is essential on MoE models and that low-SNR filtering plus tanh-modulated evidence weighting are both necessary for the final gains.
Original abstract
Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide limited guidance on which intermediate reasoning steps or tool interactions contribute to the outcome. The difficulty is especially pronounced in multi-turn search agents, where successful trajectories may contain misleading actions and failed trajectories may contain valuable evidence-gathering steps. We propose PBSD (Privileged Bayesian Self-Distillation), a Bayes-calibrated self-distillation method for fine-grained credit assignment under sparse final rewards. PBSD measures trajectory quality through the posterior-to-prior probability ratio of the verified answer and applies Bayes' rule to convert this hard-to-estimate answer-side ratio into a tractable likelihood ratio between a standard student model and a privileged answer-conditioned teacher model. Autoregressive decomposition of this Bayesian evidence score yields turn-level signals that identify whether each intermediate turn supports or undermines the verified outcome. Consequently, PBSD provides a principled and elegant reweighting scheme that transforms sparse outcome supervision into Bayes-calibrated turn-level credit signals, while remaining fully compatible with standard policy optimization. Experiments demonstrate that PBSD consistently enhances performance across both in-domain and out-of-domain settings, and effectively transfers knowledge from short-context training to long-context inference, suggesting that its fine-grained credit assignment mechanism facilitates more effective policy learning and yields improved generalization.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.