NTH

When Does Muon Help Agentic Reinforcement Learning?

AuthorsKai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun

July 25, 2026 2 min read
Watch on YouTube
The one-line take

Muon may substantially improve agentic RL training, but stronger evidence across seeds, tasks, and optimization settings is still needed.

Key results

0.546
GiGPO late-window success

Muon versus 0.290 for AdamW on ALFWorld

88%
GiGPO relative improvement

Relative late-window validation-success gain over AdamW

0.268
GRPO late-window success

Muon at 3e-5 versus 0.161 for AdamW

0.901
GraphGPO late-window success

Muon at 1e-5 versus 0.810 for AdamW

0.556
GraphGPO normalized AUC

Muon at 1e-5 versus 0.399 for AdamW

200
Training updates

Duration of each ALFWorld training run

What the paper found

This study from researchers at Renmin University of China, the Chinese Academy of Sciences, Duke University, and Zhejiang University tests whether Muon, a matrix-aware optimizer associated with large-scale training and supported in NVIDIA NeMo RL, transfers to agentic reinforcement learning. Using Qwen2.5-0.5B-Instruct on sparse-reward ALFWorld with 200 updates, the authors apply Muon only to hidden weight matrices and retain AdamW for embeddings and normalization parameters, comparing it with AdamW across GRPO, GiGPO, and GraphGPO, including GRPO’s DeepSeekMath lineage. Under GiGPO, Muon raises late-window validation success from 0.290 to 0.546, an 88% relative improvement, while matched high-rate AdamW controls collapse to zero post-update success. The effect is estimator- and learning-rate-dependent: under GRPO at 3e-5, Muon reaches 0.268 versus AdamW’s 0.161, whereas GraphGPO at 1e-5 reaches 0.901 and increases normalized validation AUC from 0.399 to 0.556, crossing 0.5 and 0.75 success 30 and 60 updates earlier. A GiGPO ablation shows Muon’s advantage persists without step-level credit, while adding step-level credit improves both optimizers. The proposed explanation is that Muon’s Newton–Schulz spectral whitening amplifies weak gradient directions only when their signs are reliable; finer-grained credit assignment may improve that signal-to-noise ratio in long-horizon tasks, unlike some negative single-turn RLVR results. The authors stress that this is exploratory evidence from one 0.5B model, one environment, and limited seeds.

Original abstract

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The effect depends on the advantage estimator and learning rate. At 3e-5, Muon improves GRPO from 0.161 to 0.268, whereas GraphGPO's late-window gap narrows near saturation. At 1e-5, GraphGPO Muon reaches 0.901, raises normalized validation AUC from 0.399 to 0.556, and reaches 0.5 and 0.75 success 30 and 60 updates earlier, respectively. These exploratory results show that Muon can benefit agentic RL and motivate studying the policy optimizer, advantage estimator, and learning rate jointly. Multi-seed and cross-task validation remain open.

Read the original paper

More in Optimization

Browse all 36 papers →