NTH

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

AuthorsAnh Nguyen, Ngan Nguyen, Duc Vu, Trung Dao, Viet Nguyen, Quan Dao, Kien Nguyen, Chi Tran, Phong Nguyen, Khoi Nguyen, Cuong Pham, Dimitris Metaxas, Vishal M. Patel, Anh Tran

July 11, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to distill powerful modern diffusion models into smaller one-step generators even when the teacher and student use incompatible latent spaces.

Key results

5M
Bridge parameters

lightweight alignment module size

1024
Teacher resolution

teacher image resolution

512
Student resolution

student image resolution

5.37
SD 1.5 HPSv3

initialized student score

9.42
SD 1.5 HPSv3 with SD 3.5 Medium

distilled student score

10.53
Merged SD 1.5 HPSv3

post-training merged checkpoint score

What the paper found

Cross-Space Distillation, from Qualcomm AI Research with collaborators at the University of Wisconsin–Madison, Johns Hopkins University, and Rutgers University, reframes one-step diffusion distillation for the real deployment case where teacher and student do not share a latent space. The paper shows that modern teachers such as Stable Diffusion 3.5, FLUX.2-klein-4B, SDXL, Kolors, and PixArt-Σ can be distilled into compact Stable Diffusion 1.5 and SD 2.1 students despite a 1024×1024 teacher-to-512×512 student resolution gap and mismatched VAE parameterizations. The key module, Bridge, is a lightweight 5M-parameter latent interface that reuses a frozen prefix of the student decoder as a spatial prior and learns a SwinIR projection head, trained with L1 latent reconstruction and an attention fidelity loss using reverse-KL on teacher self-attention. Across five teachers, the SD 1.5 student improves HPSv3 from 5.37 to 9.42 with SD 3.5 Medium, while the merged checkpoint reaches 10.53 HPSv3 and 29.07 HPSv2 without increasing the 1-step inference budget. The supplementary report shows Bridge adds only 6.01 ms to 18.01 ms runtime and 0.54 GB to 0.78 GB peak memory, and that a 30% pruned 1.5B baseline is far weaker than the 0.86B Bridge student on HPSv3, HPSv2, ImageReward, MPS, and DPG Bench. The method also supports inference-time resolution upgrade, decoding 512×512 student outputs into teacher-compatible 1024×1024 images.

Original abstract

Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inhabit the same latent space. This Shared-Space constraint prevents knowledge transfer from modern high-capacity Teachers (e.g., SD 3.5 and Flux) into compact, deployment-friendly Students such as SD 1.5, whose latent resolution and VAE parameterization differ from the Teacher. We formalize this overlooked regime as Cross-Space Distillation, where Teacher and Student differ in both latent resolution and VAE space. To enable distillation under this mismatch, we introduce the Bridge, a lightweight latent interface that maps Student latents into the Teacher space without modifying the Student backbone. Bridge combines a frozen Student VAE decoder as a spatial prior with a compact learnable projector, and is trained with latent reconstruction and attention fidelity objectives for stable Teacher-space alignment. Across diverse modern Teachers, Bridge enables substantial gains for compact one-step Students; for example, it improves SD 1.5 from 5.4 to 9.4 HPSv3 while preserving one-step inference, low latency, and broad ecosystem compatibility. These results show that heterogeneous large Teachers can be distilled into efficient, deployable backbones through a lightweight latent-space interface.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis