NTH

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR

AuthorsChongyu Fan, Gaowen Liu, Mingyi Hong, Ramana Rao Kompella, Sijia Liu

June 14, 2026 3 min read
Watch on YouTube
The one-line take

This paper introduces Pion, a faster Muon-like optimizer that keeps the useful large-gradient directions while suppressing noisy ones, boosting training for robot control and reinforcement-learning math models.

Key results

97.0%
Muon LIBERO Object success

Baseline success rate for VLA-Adapter on LIBERO Object after 1,500 steps.

32.2%
AdamW LIBERO Object success

Baseline success rate for VLA-Adapter on LIBERO Object after 1,500 steps.

85.6%
Real-robot average success

Pion on π0.5 under DROID across three grasp-and-place tasks, versus 38.9% for Muon and 31.1% for AdamW.

38.9%
Muon real-robot average success

Real-robot baseline average for Muon under DROID.

31.1%
AdamW real-robot average success

Real-robot baseline average for AdamW under DROID.

What the paper found

This paper argues that Muon, the Matrix Sign optimizer popularized in LLM pretraining, fails in two post-pretraining regimes because its uniform Newton–Schulz whitening forces every singular value toward 1. In vision-language-action training on LIBERO and LIBERO-Plus, the action head is intrinsically low-rank, so Muon amplifies tail noise instead of preserving the few informative directions; in reinforcement learning with verifiable rewards on Qwen3-1.7B and Qwen3-4B, the gradients are low-SNR, so the same whitening destabilizes policy updates and can cause collapse. The authors propose Pion, a drop-in replacement that keeps Muon’s per-step cost but replaces whitening with a two-stage high-pass Newton–Schulz filter: a Promotion polynomial with coefficients (1.875, -1.25, 0.375) followed by a Suppression polynomial with coefficients (0, 2.5, -1.5), plus an optional per-head reshape for attention projections. On VLA-Adapter, Pion reaches 100% success on LIBERO Object after 1,500 steps, versus 97.0% for Muon and 32.2% for AdamW, and on a Franka Research 3 robot under DROID it lifts average grasp-and-place success to 85.6% from 38.9% for Muon and 31.1% for AdamW. On RLVR, Muon collapses to near zero on MATH and GSM8K, while Pion restores stable training and beats AdamW across GRPO and GMPO. A reverse ablation with low-pass Muon confirms that the high-pass direction, not just the iteration form, is the critical ingredient.

Original abstract

Muon is a matrix-aware optimizer that leverages Newton-Schulz (NS) iterations to enforce spectral gradient orthogonalization by driving all singular values of the momentum matrix toward 1. While this uniform spectral whitening enhances exploration and outperforms AdamW in LLM pretraining, we show it could lead to fundamental limitations beyond pretraining in two regimes: (i) cross-modality vision-language-action (VLA) training, where inherently low-rank action-module gradients cause amplification of noisy tail directions, and (ii) reinforcement learning with verifiable rewards (RLVR), where low-SNR gradients and the need to preserve per-head specialization from prior training make whitening unstable. To address these challenges, we propose Pion, a drop-in replacement for Muon that preserves its computational efficiency while replacing uniform spectral whitening with a two-stage Promotion+Suppression mechanism, which we call the high-pass NS iteration. This design induces a sharp spectral high-pass effect, anchoring dominant singular values at 1 while suppressing noisy tail components toward 0, with controllable filter strength. To preserve pretrained per-head heterogeneity, Pion also supports a per-head mode that applies updates independently across attention heads via a simple reshape, at no extra cost. In VLA training on LIBERO and LIBERO-Plus, Pion consistently outperforms both baselines across l_1-regression (VLA-Adapter) and flow-matching (VLANeXt) architectures, e.g., reaching 100% success rate on LIBERO Object after 1,500 training steps with VLA-Adapter, vs. 97.0% for Muon and only 32.2% for AdamW. The advantage of Pion further extends to a real Franka Research 3 robot with a pi_0.5 backbone under the DROID setup on three grasp-and-place tasks. In RLVR post-training on Qwen3-1.7B/4B with GRPO and GMPO, Pion also outperforms AdamW on MATH and GSM8K while Muon collapses to zero.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →