Rethinking Cross-Layer Information Routing in Diffusion Transformers
AuthorsChao Xu, Maohua Li, Qirui Li, Yixuan Xu, Yanke Zhou, Yunhe Li, Cuifeng Shen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang
Resources
This paper argues that the way diffusion transformers pass information across layers is suboptimal, and introduces a new residual-routing method that can speed up training and improve image quality.
Key results
SiT-XL/2 on ImageNet 256×256 versus the baseline at matched compute
DAR static c4 on ImageNet 256×256 without guidance
DAR matches baseline quality in fewer training iterations
FID at 100K iterations when DAR is stacked on REPA
Fused Triton implementation versus naive forward path at N=57
What the paper found
This paper from Alibaba Group, Nanjing University, Zhejiang University, and City University of Hong Kong argues that Diffusion Transformers inherit a suboptimal residual pathway from standard Transformers, then diagnoses three depth-wise failure modes: monotonic hidden-state magnitude inflation, sharp backward gradient decay, and block-wise redundancy. On SiT-XL/2 trained on ImageNet-1K at 256×256, the authors replace residual addition with Diffusion-Adaptive Routing, or DAR, a learnable, timestep-adaptive, non-incremental softmax aggregation over historical sublayer outputs. DAR’s routing weights are conditioned either statically, dynamically through the current hidden state, or with explicit timestep injection; the timestep-aware variants are crucial because the denoising time signal is linearly decodable from router inputs with test R2 above 0.95 in early blocks and near 1.0 deeper in the stack. Empirically, DAR reaches 7.56 FID on ImageNet without guidance and 2.05 FID with classifier-free guidance, while the best reported improvement is a 2.11 FID gain over the SiT-XL/2 baseline and convergence to baseline quality in 8.75× fewer training iterations. A chunked version with chunk size 4 gives the best trade-off, and the method composes with REPA: adding DAR to REPA lowers FID from 9.89 to 7.09 at 100K iterations, showing that routing-level and representation-alignment gains are orthogonal. The paper also reports that a fused Triton implementation reduces forward latency by 11.5× and backward latency by 8.5× at the SiT-XL/2 working point.
Original abstract
Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively revisited. The residual stream that governs how information accumulates across layers, however, has been directly inherited from the original Transformer. In this paper, we present a systematic empirical analysis of cross-layer information flow in DiTs, jointly along depth and denoising timestep, and identify three concrete symptoms of traditional residual addition, namely monotonic forward magnitude inflation, sharp backward gradient decay, and pronounced block-wise redundancy. Motivated by this diagnosis, we propose Diffusion-Adaptive Routing (\textsc{DAR}), a drop-in residual replacement that performs \emph{learnable, timestep-adaptive, and non-incremental} aggregation over the history of sublayer outputs. Moreover, the proposed \textsc{DAR} is compatible with many modern Transformer enhancement methods, such as REPA. On ImageNet $256\times256$, \textsc{DAR} improves SiT-XL/2 by $2.11$ FID ($7.56$ vs.\ $9.67$) and matches the baseline's converged quality with $8.75\times$ fewer training iterations. Stacked on top of REPA, it yields a $2\times$ training acceleration in the early stage, suggesting cross-layer information routing as an underexplored design axis in diffusion modeling, one that operates orthogonally to existing representation-alignment objectives. Beyond pretraining, \textsc{DAR} can also be applied during the fine-tuning stage of large-scale T2I models and preserves high-frequency details during Distribution Matching Distillation.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.