NTH

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

AuthorsChuyan Chen, Peng Sun, Kun Yuan

August 7, 2026 2 min read
Watch on YouTube
The one-line take

CMuon makes diffusion-transformer training faster and more stable by orthogonalizing fused weight components independently instead of coupling them together.

Key results

675M
DiT-XL model size

Parameter count of the ImageNet-1K DiT-XL model

1.18
CMuon FID

FID on ImageNet-1K at 256 × 256 after 200 epochs

2×
Training speedup

CMuon speedup over AdamW

3.02
Joint chunking ablation FID

FID for 130M DiT-B after 200 epochs with FFN, QKV, and AdaLN chunking

9.07
CMuon throughput

Iterations per second for DiT-B on 8 H100 GPUs with system optimizations

What the paper found

Researchers Chuyan Chen, Peng Sun, and Kun Yuan from Peking University, Westlake University, and Zhejiang University introduce CMuon, an optimizer designed to fix Muon’s late-stage convergence plateau when training Diffusion Transformers. Standard DiT implementations fuse functionally separate weights in AdaLN modulation, attention QKV projections, and FFN gate-up layers; applying Newton–Schulz momentum orthogonalization to each fused matrix creates unwanted cross-subspace coupling and distorts update directions. CMuon reverses this fusion logically by splitting selected matrices into independent chunks, orthogonalizing each chunk separately, and preserving the overall update norm with Moonlight scaling. On ImageNet-1K at 256 × 256, a 675M-parameter DiT-XL using VA-VAE reaches FID 1.18 after 200 epochs, compared with AdamW’s FID 1.21 after 400 epochs, delivering a 2× training speedup without auxiliary methods such as REPA. On the 130M DiT-B ablation, jointly chunking FFN, QKV, and AdaLN lowers FID from 3.31 to 3.02 at 200 epochs. System optimizations, including batched Newton–Schulz kernels and overlapped communication, raise CMuon throughput to 9.07 iterations per second for DiT-B on 8 H100 GPUs, close to Adam’s 9.51, while adding negligible algorithmic overhead.

Original abstract

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence. In this paper, we identify the root cause of this bottleneck: standard DiT architectures fuse functionally distinct weights (e.g., within AdaLN and QKV layers) into unified tensors for computational efficiency. Applying Muon to these fused tensors inadvertently induces implicit subspace coupling, which distorts update directions and degrades global optimization. To address this, we introduce Chunked Muon (CMuon), a simple yet highly effective strategy that partitions these matrices into independent sub-components prior to orthogonalization. Extensive experiments demonstrate that a 675M-parameter DiT trained with CMuon achieves a FID of 1.18 on ImageNet 256 in just 200 epochs. This represents more than a 2x training speedup over AdamW, while effectively overcoming the late-stage convergence plateaus of vanilla Muon.

Read the original paper

More in Optimization

Browse all 36 papers →