Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
AuthorsHaoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang, Yu Liu, Changxin Gao, Nong Sang
Resources
SharpMoE improves diffusion model efficiency and image quality by routing computation to the most important tokens using cleaner saliency signals instead of noisy latents.
Key results
Evaluation setup for class-conditional generation
SharpMoE adaptation length on pretrained backbones
Initialization checkpoint before SharpMoE post-training
Recursive training horizon used in SharpMoE
Balance between flow matching and routing alignment
Best reported FID50K for SharpMoE at CFG 1.5
What the paper found
SharpMoE, developed by researchers from Huazhong University of Science and Technology and Alibaba Group’s Tongyi Lab, targets a specific failure mode in diffusion Mixture-of-Experts models: routers trained on noisy latents do not reliably allocate more experts to salient image tokens. The paper’s core idea is to replace noisy routing with saliency-harnessed clean routing by feeding the router the previous step’s predicted clean latent x̂0, which preserves structural and textural cues even at high-noise timesteps. SharpMoE is implemented as a post-training, plug-and-play enhancement on pretrained Diffusion Transformer backbones, including TC-DiT, EC-DiT, and DiffMoE, and is trained with a recursive full-trajectory scheme over T = 10 rollout steps. To explicitly align compute with visual importance, the authors add a trajectory routing loss that matches cumulative expert assignment to a Laplacian-based saliency map using KL divergence, with λrouting = 0.001. On ImageNet 256×256, after 100K post-training steps from 500K-pretrained models, SharpMoE improves every evaluated backbone; for example, DiffMoE-L reaches FID50K 3.10 and IS 228.88 at CFG 1.5, while the baseline DiffMoE-L scores 3.86 and 203.00. Ablations show that saliency-harnessing routing reduces DiffMoE-B FID50K from 8.03 to 6.95 at CFG 1.5, and adding the trajectory routing loss further lowers it to 6.66, confirming that clean latent guidance and trajectory-level compute regularization are both necessary for accurate saliency-aware expert allocation.
Original abstract
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocating computational resources across diverse tokens to improve efficiency and performance. However, we identify a routing assignment problem in existing diffusion MoE frameworks: the router fails to accurately allocate more computational resources to salient tokens. Our analysis attributes this failure to the router's reliance on noise-corrupted latent features throughout the denoising process. Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. By bypassing the noise-distorted inputs, SharpMoE provides the router with clear saliency guidance, enabling the identification of salient tokens even in high-noise stages. Furthermore, we introduce a trajectory routing loss to constrain the compute allocation throughout the multi-step denoising trajectory, ensuring precise resource allocation along the generation rollout. Extensive experiments demonstrate that SharpMoE serves as a versatile, plug-and-play solution that further enhances the pretrained, converged MoE models, achieving state-of-the-art performance in visual generation.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.