NTH

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

AuthorsHaoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang, Yu Liu, Changxin Gao, Nong Sang

July 2, 2026 2 min read
Watch on YouTube
The one-line take

SharpMoE improves diffusion model efficiency and image quality by routing computation to the most important tokens using cleaner saliency signals instead of noisy latents.

Key results

256×256
ImageNet resolution

Evaluation setup for class-conditional generation

100K
Post-training steps

SharpMoE adaptation length on pretrained backbones

500K
Pretraining steps

Initialization checkpoint before SharpMoE post-training

10
Full-trajectory rollout steps

Recursive training horizon used in SharpMoE

0.001
Trajectory routing loss weight

Balance between flow matching and routing alignment

3.10
DiffMoE-L FID50K

Best reported FID50K for SharpMoE at CFG 1.5

What the paper found

SharpMoE, developed by researchers from Huazhong University of Science and Technology and Alibaba Group’s Tongyi Lab, targets a specific failure mode in diffusion Mixture-of-Experts models: routers trained on noisy latents do not reliably allocate more experts to salient image tokens. The paper’s core idea is to replace noisy routing with saliency-harnessed clean routing by feeding the router the previous step’s predicted clean latent x̂0, which preserves structural and textural cues even at high-noise timesteps. SharpMoE is implemented as a post-training, plug-and-play enhancement on pretrained Diffusion Transformer backbones, including TC-DiT, EC-DiT, and DiffMoE, and is trained with a recursive full-trajectory scheme over T = 10 rollout steps. To explicitly align compute with visual importance, the authors add a trajectory routing loss that matches cumulative expert assignment to a Laplacian-based saliency map using KL divergence, with λrouting = 0.001. On ImageNet 256×256, after 100K post-training steps from 500K-pretrained models, SharpMoE improves every evaluated backbone; for example, DiffMoE-L reaches FID50K 3.10 and IS 228.88 at CFG 1.5, while the baseline DiffMoE-L scores 3.86 and 203.00. Ablations show that saliency-harnessing routing reduces DiffMoE-B FID50K from 8.03 to 6.95 at CFG 1.5, and adding the trajectory routing loss further lowers it to 6.66, confirming that clean latent guidance and trajectory-level compute regularization are both necessary for accurate saliency-aware expert allocation.

Original abstract

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocating computational resources across diverse tokens to improve efficiency and performance. However, we identify a routing assignment problem in existing diffusion MoE frameworks: the router fails to accurately allocate more computational resources to salient tokens. Our analysis attributes this failure to the router's reliance on noise-corrupted latent features throughout the denoising process. Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. By bypassing the noise-distorted inputs, SharpMoE provides the router with clear saliency guidance, enabling the identification of salient tokens even in high-noise stages. Furthermore, we introduce a trajectory routing loss to constrain the compute allocation throughout the multi-step denoising trajectory, ensuring precise resource allocation along the generation rollout. Extensive experiments demonstrate that SharpMoE serves as a versatile, plug-and-play solution that further enhances the pretrained, converged MoE models, achieving state-of-the-art performance in visual generation.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis