When Does Muon Help Agentic Reinforcement Learning?
AuthorsKai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
Resources
Muon may substantially improve agentic RL training, but stronger evidence across seeds, tasks, and optimization settings is still needed.
Key results
Muon versus 0.290 for AdamW on ALFWorld
Relative late-window validation-success gain over AdamW
Muon at 3e-5 versus 0.161 for AdamW
Muon at 1e-5 versus 0.810 for AdamW
Muon at 1e-5 versus 0.399 for AdamW
Duration of each ALFWorld training run
What the paper found
This study from researchers at Renmin University of China, the Chinese Academy of Sciences, Duke University, and Zhejiang University tests whether Muon, a matrix-aware optimizer associated with large-scale training and supported in NVIDIA NeMo RL, transfers to agentic reinforcement learning. Using Qwen2.5-0.5B-Instruct on sparse-reward ALFWorld with 200 updates, the authors apply Muon only to hidden weight matrices and retain AdamW for embeddings and normalization parameters, comparing it with AdamW across GRPO, GiGPO, and GraphGPO, including GRPO’s DeepSeekMath lineage. Under GiGPO, Muon raises late-window validation success from 0.290 to 0.546, an 88% relative improvement, while matched high-rate AdamW controls collapse to zero post-update success. The effect is estimator- and learning-rate-dependent: under GRPO at 3e-5, Muon reaches 0.268 versus AdamW’s 0.161, whereas GraphGPO at 1e-5 reaches 0.901 and increases normalized validation AUC from 0.399 to 0.556, crossing 0.5 and 0.75 success 30 and 60 updates earlier. A GiGPO ablation shows Muon’s advantage persists without step-level credit, while adding step-level credit improves both optimizers. The proposed explanation is that Muon’s Newton–Schulz spectral whitening amplifies weak gradient directions only when their signs are reliable; finer-grained credit assignment may improve that signal-to-noise ratio in long-horizon tasks, unlike some negative single-turn RLVR results. The authors stress that this is exploratory evidence from one 0.5B model, one environment, and limited seeds.
Original abstract
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The effect depends on the advantage estimator and learning rate. At 3e-5, Muon improves GRPO from 0.161 to 0.268, whereas GraphGPO's late-window gap narrows near saturation. At 1e-5, GraphGPO Muon reaches 0.901, raises normalized validation AUC from 0.399 to 0.556, and reaches 0.5 and 0.75 success 30 and 60 updates earlier, respectively. These exploratory results show that Muon can benefit agentic RL and motivate studying the policy optimizer, advantage estimator, and learning rate jointly. Multi-seed and cross-task validation remain open.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.