NTH

RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models

AuthorsXing Cong, Hanlin Tang, Kan Liu, Lan Tao, Lin Qu, Chenhao Xie

June 9, 2026 2 min read
Watch on YouTube
The one-line take

RT-Lynx speeds up diffusion transformers by sparsifying activations instead of weights, preserving image quality while making generation faster.

Key results

5-10%
active neuron ratio

Fraction of neurons activated in token-wise DiT activations

64
LoRA rank

Low-rank compensation setting used in RT-Lynx

21.25
MJHQ FID

RT-Lynx result on Qwen-Image

1.304
MJHQ Image Reward

RT-Lynx result on Qwen-Image

66.91
sparse-weight FID

Naive 2:4 weight sparsity on Qwen-Image over sDCI

1.88x
linear-layer speedup

Maximum speedup reported for the optimized sparse GEMM kernel

What the paper found

RT-Lynx, from Alibaba Group researchers, argues that Diffusion Transformers should be sparsified in activations rather than weights: across Qwen-Image, FLUX.1-dev, and Z-Image, token-wise activations are intrinsically sparse, with only about 5% to 10% of neurons active, while weight sparsification under the 2:4 pattern destroys generation quality. The method applies online N:M sparsification to intermediate activations, rescales them with norm compensation, and restores residual detail through a lightweight LoRA branch with rank 64; for single-stream models it also skips the most sensitive mixed layers. On Qwen-Image, RT-Lynx reaches FID 21.25 and Image Reward 1.304 on MJHQ-30K, improving over the dense baseline FID 21.98 and outperforming prior weight-sparsification methods such as Wanda, RIA, BaWA, and Slim. It also stays strong on sDCI with FID 25.78 and Image Reward 1.226, while the raw sparse-weight baseline collapses to FID 66.91. System-wise, a fused CUDA pipeline for online activation sparsification and sparse Tensor Core execution reduces sparse overhead to as low as 2.28% on H20 and delivers up to 1.88× speedup in linear layers and about 1.2× end-to-end acceleration, cutting Qwen-Image generation time from 0.75 s to 0.62 s without noticeable quality loss.

Original abstract

Diffusion Transformers (DiT) achieve strong performance in image generation but incur substantial inference costs. While prior work has reduced this cost via quantization and distillation, semi-structured sparsity, which can nearly halve FLOPs, remains underexplored. A key reason is that most existing approaches focus on weight sparsification, and pruning 50% of the weights can remove critical model capacity and degrade generation quality. Our study, however, shows that DiT activations are intrinsically sparse and significantly more robust to N:M semi-structured sparsification than weights. Motivated by this observation, we advocate a paradigm shift from weight sparsification to activation sparsification. We propose RT-Lynx, which applies N:M sparsification to activations and incorporates error-compensation techniques to mitigate accuracy loss. We further implement highly optimized CUDA kernels tailored to this setting, achieving up to a 1.55x speedup on average in linear layers. Extensive experiments across multiple diffusion models demonstrate that our method preserves the generation quality of the original models while substantially accelerating inference.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis