NTH

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

AuthorsYang Tian, Rui Wang, Xumeng Wen, Junjie Li, Shizhao Sun, Lei Song, Jiang Bian, Bo Zhao

June 18, 2026 2 min read
Watch on YouTube
The one-line take

PBSD turns sparse final rewards into step-by-step learning signals for agents by using Bayesian self-distillation to figure out which actions really helped.

Key results

40.87
Validation score

PBSD in-domain validation accuracy

35.83
BC(300) score

PBSD performance on the stratified BrowseComp subset

2.62
GRPO validation gain

PBSD improvement over GRPO on validation

3.50
GRPO BC(300) gain

PBSD improvement over GRPO on BC(300)

46.21
SFT+PBSD BrowseComp

Best trained-agent BrowseComp score reported in the benchmark table

What the paper found

PBSD, or Privileged Bayesian Self-Distillation, is a long-horizon credit-assignment method for reinforcement learning with verifiable rewards that targets the core weakness of outcome-only training: a single final reward cannot tell which of hundreds of search and tool-use turns were actually useful. The paper, from Shanghai Jiao Tong University and XYZ AI Lab, reformulates trajectory quality as a posterior-to-prior evidence ratio for the verified answer, then uses Bayes’ rule to convert that intractable quantity into a turn-level likelihood ratio between a standard student policy and an answer-conditioned privileged teacher. This Bayesian evidence score is autoregressively decomposed into per-turn signals and used only as a detached weight on GRPO-style advantages, avoiding direct imitation and information leakage. The method is evaluated on Qwen3-30B-A3B-Thinking-2507 with 64K training context and up to 300 interaction turns, using 7.5K supervised trajectories and 575 RL training examples from a synthetic Wikipedia/OpenSeeker pipeline. On the in-domain validation set and a stratified 300-example BrowseComp subset, PBSD outperforms GRPO by 2.62 and 3.50 points, reaching 40.87 and 35.83 respectively, and it sets the best trained-agent BrowseComp score in the broader benchmark table at 46.21 with SFT+PBSD. Ablations show that replay-free evidence scoring is essential on MoE models and that low-SNR filtering plus tanh-modulated evidence weighting are both necessary for the final gains.

Original abstract

Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide limited guidance on which intermediate reasoning steps or tool interactions contribute to the outcome. The difficulty is especially pronounced in multi-turn search agents, where successful trajectories may contain misleading actions and failed trajectories may contain valuable evidence-gathering steps. We propose PBSD (Privileged Bayesian Self-Distillation), a Bayes-calibrated self-distillation method for fine-grained credit assignment under sparse final rewards. PBSD measures trajectory quality through the posterior-to-prior probability ratio of the verified answer and applies Bayes' rule to convert this hard-to-estimate answer-side ratio into a tractable likelihood ratio between a standard student model and a privileged answer-conditioned teacher model. Autoregressive decomposition of this Bayesian evidence score yields turn-level signals that identify whether each intermediate turn supports or undermines the verified outcome. Consequently, PBSD provides a principled and elegant reweighting scheme that transforms sparse outcome supervision into Bayes-calibrated turn-level credit signals, while remaining fully compatible with standard policy optimization. Experiments demonstrate that PBSD consistently enhances performance across both in-domain and out-of-domain settings, and effectively transfers knowledge from short-context training to long-context inference, suggesting that its fine-grained credit assignment mechanism facilitates more effective policy learning and yields improved generalization.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →