CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
AuthorsChuyan Chen, Peng Sun, Kun Yuan
Resources
CMuon makes diffusion-transformer training faster and more stable by orthogonalizing fused weight components independently instead of coupling them together.
Key results
Parameter count of the ImageNet-1K DiT-XL model
FID on ImageNet-1K at 256 × 256 after 200 epochs
CMuon speedup over AdamW
FID for 130M DiT-B after 200 epochs with FFN, QKV, and AdaLN chunking
Iterations per second for DiT-B on 8 H100 GPUs with system optimizations
What the paper found
Researchers Chuyan Chen, Peng Sun, and Kun Yuan from Peking University, Westlake University, and Zhejiang University introduce CMuon, an optimizer designed to fix Muon’s late-stage convergence plateau when training Diffusion Transformers. Standard DiT implementations fuse functionally separate weights in AdaLN modulation, attention QKV projections, and FFN gate-up layers; applying Newton–Schulz momentum orthogonalization to each fused matrix creates unwanted cross-subspace coupling and distorts update directions. CMuon reverses this fusion logically by splitting selected matrices into independent chunks, orthogonalizing each chunk separately, and preserving the overall update norm with Moonlight scaling. On ImageNet-1K at 256 × 256, a 675M-parameter DiT-XL using VA-VAE reaches FID 1.18 after 200 epochs, compared with AdamW’s FID 1.21 after 400 epochs, delivering a 2× training speedup without auxiliary methods such as REPA. On the 130M DiT-B ablation, jointly chunking FFN, QKV, and AdaLN lowers FID from 3.31 to 3.02 at 200 epochs. System optimizations, including batched Newton–Schulz kernels and overlapped communication, raise CMuon throughput to 9.07 iterations per second for DiT-B on 8 H100 GPUs, close to Adam’s 9.51, while adding negligible algorithmic overhead.
Original abstract
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence. In this paper, we identify the root cause of this bottleneck: standard DiT architectures fuse functionally distinct weights (e.g., within AdaLN and QKV layers) into unified tensors for computational efficiency. Applying Muon to these fused tensors inadvertently induces implicit subspace coupling, which distorts update directions and degrades global optimization. To address this, we introduce Chunked Muon (CMuon), a simple yet highly effective strategy that partitions these matrices into independent sub-components prior to orthogonalization. Extensive experiments demonstrate that a 675M-parameter DiT trained with CMuon achieves a FID of 1.18 on ImageNet 256 in just 200 epochs. This represents more than a 2x training speedup over AdamW, while effectively overcoming the late-stage convergence plateaus of vanilla Muon.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.