Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
AuthorsXingjian Wu, Junlin Liu, Xingchen Liu, Xuhang Zhu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai
Resources
CRPO combines contrastive learning with reinforcement-based self-distillation to make agentic LLM training more stable, exploratory, and generalizable.
Key results
Number of reasoning and deep-search benchmarks used for evaluation
Default fraction of entropy-ranked positions treated as positive pairs
Default λ controlling the CRPO regularizer in CRPO*
Qwen3-14B CRPO* result
Qwen3-14B CRPO* result
Qwen3-14B CRPO* result
What the paper found
Contrastive Reinforced Policy Optimization, or CRPO, addresses a failure mode in on-policy self-distillation for long-horizon language-model agents: privileged self-teachers can become overconfident after tool calls, copying demonstrated reasoning routes instead of supporting generalization. CRPO treats the original and feedback-augmented contexts as two views, uses predictive-entropy differences to classify token positions into reflective-exploration positives and exposure-biased negatives, and applies group-wise InfoNCE over negative KL similarity. Positive positions are pulled toward the self-teacher, while negative positions are explicitly repelled, producing a contrastively reweighted token-level policy gradient. The standalone CRPO objective can also replace GRPO’s static reference-model regularizer in CRPO*, preserving outcome-level reinforcement learning while adding dense, position-aware supervision. Across 13 reasoning and deep-search benchmarks, experiments with Qwen2.5, Llama3.1, Qwen3, and comparisons against DeepSeek-R1 show that CRPO* is strongest when the positive-pair proportion is 30% and the contrastive weight is λ=5. On Qwen3-14B, CRPO* reaches 70.7% Pass@5 on GAIA, 35.0% on Humanity’s Last Exam, 68.9% on WebWalkerQA, and 60.7% on xbench, demonstrating improved sampling diversity and long-horizon tool-use performance without relying solely on sparse outcome rewards.
Original abstract
Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions. To tackle these challenges, we introduce Contrastive Reinforced Policy Optimization (CRPO), which reformulates agentic OPSD from a contrastive learning perspective. By leveraging predictive entropy to distinguish between positive positions (reflective exploration) and negative positions (exposure bias), CRPO conducts group-wise contrast to preserve reliable, fine-grained optimization signals. Extensive evaluations across 13 challenging reasoning and deep-search benchmarks demonstrate that CRPO consistently outperforms existing reinforcement learning and self-distillation baselines, significantly enhancing training stability and generalization in long-horizon interactions.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.