NTH

Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization

AuthorsHao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu, Xin Liu, Qihua Chen, Hongtao Duan, Lulu Hu, Bin Yang, Minying Zhang

June 14, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces an information-bottleneck-guided way to balance exploration and exploitation in LLM reinforcement learning, boosting reasoning performance with smarter tree-based sampling.

Key results

17K
DAPO-Math-17K

training set used for online RL

50%
trajectory gain

IBTree samples more trajectories under the same token budget

29.2
overall score Qwen3-1.7B

avg@32 overall benchmark score for IB-TPO

44.3
overall score Qwen3-8B

avg@32 overall benchmark score for IB-TPO

What the paper found

This Alibaba Cloud Computing paper introduces IB-Score, a step-level metric derived from Information Bottleneck theory that quantifies the exploration–exploitation balance in online reinforcement learning for large language models by coupling reasoning diversity with mutual information to the correct answer. The authors show that standard GRPO training on Qwen3-8B-Base quickly loses covariance between information gain and confidence, collapsing into over-exploitation, while entropy regularization can instead trigger over-exploration and instability. To fix this, they propose IB-TPO, which combines an IB-guided tree sampling method called IBTree with an IB-based advantage objective, reusing the tree structure for Monte Carlo estimation and producing 50% more trajectories under the same token budget. On DAPO-Math-17K, evaluated with avg@32 across MATH-500, AIME 24/25, AMC 23/24, GPQA Diamond, and IFEval, IB-TPO reaches 29.2 overall on Qwen3-1.7B-Base versus 26.3 for vanilla GRPO, and 44.3 versus 40.7 on Qwen3-8B-Base, outperforming TreeRL, TreePO, and IBRO. Ablations show the full IBTree plus IBTPO advantage formulation is best, and the method remains strong at 4K and 8K context lengths; the authors note the main remaining cost is multi-iteration tree sampling latency.

Original abstract

Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning tasks. However, they often exhibit an imbalanced exploration-exploitation trade-off, resulting in unstable optimization and sub-optimal performance. We introduce IB-Score, a novel metric grounded in Information Bottleneck theory that evaluates policy's exploration-exploitation balance by quantifying the trade-off between step-level reasoning diversity and mutual information shared with the correct answer. Analysis based on IB-Score shows that popular online RL approaches (e.g., GRPO) with common regularizers fail to consistently maintain balance during training with suboptimal results. To address this, we propose Information Bottleneck-driven Tree-based Policy Optimization (IB-TPO), a principled framework that formulates IB-Score as a fine-grained optimization objective and utilizes a novel IB-guided tree sampling strategy that not only improves the efficiency of online sampling with 50% more trajectories under the same token budget, but also reuses the tree structure for effective IB-Score Monte Carlo estimation. Extensive experiments across standard benchmarks show that our method significantly outperforms GRPO baseline by 2.9% to 3.6% and also outperforms other state-of-the-art online RL approaches. Our code is available at https://github.com/alibaba/EfficientRL.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →