Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
AuthorsXianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu
ERPO regularizes the queries an LLM trains on instead of directly constraining its responses, aiming for more stable learning without sacrificing exploration.
Key results
Overall improvement over GRPO across six mathematical reasoning benchmarks.
Overall Pass@1 improvement over GRPO.
ERPO score, compared with 0.528 for GRPO.
ERPO reduces the average gap from 6.47% to 3.14%.
Accuracy from temperatures 1.2 to 1.5, compared with 57.2% for GRPO.
What the paper found
This paper introduces Environment-Regularized Policy Optimization, or ERPO, to address a central failure mode in LLM reinforcement learning: standard action-side Policy-KL can stabilize response updates but does not prevent the model-induced query distribution from drifting. ERPO replaces Policy-KL with Query-KL, which regularizes the autoregressive likelihood of training prompts against a pre-RL reference while sending gradients only through query likelihood, leaving response-side exploration unconstrained. It also applies a cached, reference-derived per-query weight to reduce gradient variance and favor prompts typical of the original model, with no additional forward passes. The method is a drop-in extension for GRPO, PPO, and REINFORCE. Experiments train Qwen2.5-Math-7B and Qwen2.5-32B using EasyR1 on approximately 8.5K Level 3–5 MATH examples, then evaluate AIME24, AIME25, AMC, MATH500, Minerva, and OlympiadBench across sampling temperatures from 0.1 to 1.5. Against GRPO, ERPO improves overall Avg@32 by 6.2%, Pass@32 by 3.64%, and Pass@1 by 5.69%; on MATH500, mean Avg@32 reaches 0.677 versus 0.528 for GRPO. ERPO also reduces the average train–evaluation gap from 6.47% to 3.14%, a 51% reduction, mitigating reward hacking during long-horizon training. On Qwen2.5-32B, ERPO reaches 82.8% accuracy at high temperatures from 1.2 to 1.5, compared with 57.2% for GRPO.
Original abstract
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in a double bind: keeping Policy-KL constrains response behavior and consumes the action-side exploration budget, while dropping it leaves the optimization without an explicit drift control. We argue for an alternative that breaks the dilemma by moving regularization to the input side. As training progresses, the distribution over training queries induced by the current policy drifts unchecked from its pre-RL reference distribution. Concretely, Environment-Regularized Policy Optimization (ERPO) introduces a Query-KL (QKL) term that bounds this query distribution shift, together with a dataset-static reference-derived per-query weight that biases each per-query update toward queries typical under the reference. The QKL gradient flows strictly through the query likelihood; the response score function used by policy-gradient estimators does not appear in the QKL term, so QKL exerts no direct gradient pressure on the response distribution---exploration is preserved. ERPO plugs into GRPO/PPO/REINFORCE-style pipelines without additional forward passes. On six mathematical reasoning benchmarks, ERPO replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.Our source code are available at https://github.com/alibaba/ERPO
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.