NTH

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

AuthorsXingjian Wu, Junlin Liu, Xingchen Liu, Xuhang Zhu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai

August 9, 2026 2 min read
Watch on YouTube
The one-line take

CRPO combines contrastive learning with reinforcement-based self-distillation to make agentic LLM training more stable, exploratory, and generalizable.

Key results

13
Benchmark coverage

Number of reasoning and deep-search benchmarks used for evaluation

30%
Positive-pair proportion

Default fraction of entropy-ranked positions treated as positive pairs

5
Contrastive weight

Default λ controlling the CRPO regularizer in CRPO*

70.7%
GAIA Pass@5

Qwen3-14B CRPO* result

35.0%
Humanity’s Last Exam Pass@5

Qwen3-14B CRPO* result

68.9%
WebWalkerQA Pass@5

Qwen3-14B CRPO* result

What the paper found

Contrastive Reinforced Policy Optimization, or CRPO, addresses a failure mode in on-policy self-distillation for long-horizon language-model agents: privileged self-teachers can become overconfident after tool calls, copying demonstrated reasoning routes instead of supporting generalization. CRPO treats the original and feedback-augmented contexts as two views, uses predictive-entropy differences to classify token positions into reflective-exploration positives and exposure-biased negatives, and applies group-wise InfoNCE over negative KL similarity. Positive positions are pulled toward the self-teacher, while negative positions are explicitly repelled, producing a contrastively reweighted token-level policy gradient. The standalone CRPO objective can also replace GRPO’s static reference-model regularizer in CRPO*, preserving outcome-level reinforcement learning while adding dense, position-aware supervision. Across 13 reasoning and deep-search benchmarks, experiments with Qwen2.5, Llama3.1, Qwen3, and comparisons against DeepSeek-R1 show that CRPO* is strongest when the positive-pair proportion is 30% and the contrastive weight is λ=5. On Qwen3-14B, CRPO* reaches 70.7% Pass@5 on GAIA, 35.0% on Humanity’s Last Exam, 68.9% on WebWalkerQA, and 60.7% on xbench, demonstrating improved sampling diversity and long-horizon tool-use performance without relying solely on sparse outcome rewards.

Original abstract

Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions. To tackle these challenges, we introduce Contrastive Reinforced Policy Optimization (CRPO), which reformulates agentic OPSD from a contrastive learning perspective. By leveraging predictive entropy to distinguish between positive positions (reflective exploration) and negative positions (exposure bias), CRPO conducts group-wise contrast to preserve reliable, fine-grained optimization signals. Extensive evaluations across 13 challenging reasoning and deep-search benchmarks demonstrate that CRPO consistently outperforms existing reinforcement learning and self-distillation baselines, significantly enhancing training stability and generalization in long-horizon interactions.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →