Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization
AuthorsHao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu, Xin Liu, Qihua Chen, Hongtao Duan, Lulu Hu, Bin Yang, Minying Zhang
This paper introduces an information-bottleneck-guided way to balance exploration and exploitation in LLM reinforcement learning, boosting reasoning performance with smarter tree-based sampling.
Key results
training set used for online RL
IBTree samples more trajectories under the same token budget
avg@32 overall benchmark score for IB-TPO
avg@32 overall benchmark score for IB-TPO
What the paper found
This Alibaba Cloud Computing paper introduces IB-Score, a step-level metric derived from Information Bottleneck theory that quantifies the exploration–exploitation balance in online reinforcement learning for large language models by coupling reasoning diversity with mutual information to the correct answer. The authors show that standard GRPO training on Qwen3-8B-Base quickly loses covariance between information gain and confidence, collapsing into over-exploitation, while entropy regularization can instead trigger over-exploration and instability. To fix this, they propose IB-TPO, which combines an IB-guided tree sampling method called IBTree with an IB-based advantage objective, reusing the tree structure for Monte Carlo estimation and producing 50% more trajectories under the same token budget. On DAPO-Math-17K, evaluated with avg@32 across MATH-500, AIME 24/25, AMC 23/24, GPQA Diamond, and IFEval, IB-TPO reaches 29.2 overall on Qwen3-1.7B-Base versus 26.3 for vanilla GRPO, and 44.3 versus 40.7 on Qwen3-8B-Base, outperforming TreeRL, TreePO, and IBRO. Ablations show the full IBTree plus IBTPO advantage formulation is best, and the method remains strong at 4K and 8K context lengths; the authors note the main remaining cost is multi-iteration tree sampling latency.
Original abstract
Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning tasks. However, they often exhibit an imbalanced exploration-exploitation trade-off, resulting in unstable optimization and sub-optimal performance. We introduce IB-Score, a novel metric grounded in Information Bottleneck theory that evaluates policy's exploration-exploitation balance by quantifying the trade-off between step-level reasoning diversity and mutual information shared with the correct answer. Analysis based on IB-Score shows that popular online RL approaches (e.g., GRPO) with common regularizers fail to consistently maintain balance during training with suboptimal results. To address this, we propose Information Bottleneck-driven Tree-based Policy Optimization (IB-TPO), a principled framework that formulates IB-Score as a fine-grained optimization objective and utilizes a novel IB-guided tree sampling strategy that not only improves the efficiency of online sampling with 50% more trajectories under the same token budget, but also reuses the tree structure for effective IB-Score Monte Carlo estimation. Extensive experiments across standard benchmarks show that our method significantly outperforms GRPO baseline by 2.9% to 3.6% and also outperforms other state-of-the-art online RL approaches. Our code is available at https://github.com/alibaba/EfficientRL.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.