Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
AuthorsAnh Nguyen, Ngan Nguyen, Duc Vu, Trung Dao, Viet Nguyen, Quan Dao, Kien Nguyen, Chi Tran, Phong Nguyen, Khoi Nguyen, Cuong Pham, Dimitris Metaxas, Vishal M. Patel, Anh Tran
Resources
This paper shows how to distill powerful modern diffusion models into smaller one-step generators even when the teacher and student use incompatible latent spaces.
Key results
lightweight alignment module size
teacher image resolution
student image resolution
initialized student score
distilled student score
post-training merged checkpoint score
What the paper found
Cross-Space Distillation, from Qualcomm AI Research with collaborators at the University of Wisconsin–Madison, Johns Hopkins University, and Rutgers University, reframes one-step diffusion distillation for the real deployment case where teacher and student do not share a latent space. The paper shows that modern teachers such as Stable Diffusion 3.5, FLUX.2-klein-4B, SDXL, Kolors, and PixArt-Σ can be distilled into compact Stable Diffusion 1.5 and SD 2.1 students despite a 1024×1024 teacher-to-512×512 student resolution gap and mismatched VAE parameterizations. The key module, Bridge, is a lightweight 5M-parameter latent interface that reuses a frozen prefix of the student decoder as a spatial prior and learns a SwinIR projection head, trained with L1 latent reconstruction and an attention fidelity loss using reverse-KL on teacher self-attention. Across five teachers, the SD 1.5 student improves HPSv3 from 5.37 to 9.42 with SD 3.5 Medium, while the merged checkpoint reaches 10.53 HPSv3 and 29.07 HPSv2 without increasing the 1-step inference budget. The supplementary report shows Bridge adds only 6.01 ms to 18.01 ms runtime and 0.54 GB to 0.78 GB peak memory, and that a 30% pruned 1.5B baseline is far weaker than the 0.86B Bridge student on HPSv3, HPSv2, ImageReward, MPS, and DPG Bench. The method also supports inference-time resolution upgrade, decoding 512×512 student outputs into teacher-compatible 1024×1024 images.
Original abstract
Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inhabit the same latent space. This Shared-Space constraint prevents knowledge transfer from modern high-capacity Teachers (e.g., SD 3.5 and Flux) into compact, deployment-friendly Students such as SD 1.5, whose latent resolution and VAE parameterization differ from the Teacher. We formalize this overlooked regime as Cross-Space Distillation, where Teacher and Student differ in both latent resolution and VAE space. To enable distillation under this mismatch, we introduce the Bridge, a lightweight latent interface that maps Student latents into the Teacher space without modifying the Student backbone. Bridge combines a frozen Student VAE decoder as a spatial prior with a compact learnable projector, and is trained with latent reconstruction and attention fidelity objectives for stable Teacher-space alignment. Across diverse modern Teachers, Bridge enables substantial gains for compact one-step Students; for example, it improves SD 1.5 from 5.4 to 9.4 HPSv3 while preserving one-step inference, low latency, and broad ecosystem compatibility. These results show that heterogeneous large Teachers can be distilled into efficient, deployable backbones through a lightweight latent-space interface.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.