PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
AuthorsYifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang, Huan Ling, Sanja Fidler, Xuanchi Ren
Resources
PiD replaces traditional latent decoders with a fast pixel-diffusion decoder that can turn compressed latents into high-resolution images much more efficiently and with better visual fidelity.
Key results
High-quality filtered dataset used to train PiD
PixelDiT checkpoint used as the high-resolution pixel diffusion prior
DMD2-distilled decoder step count
Peak GPU memory in GB for 2048×2048 decoding
End-to-end decoding time in ms for 2048×2048 outputs
Approximate speedup over diffusion-based super-resolution pipelines
What the paper found
PiD, from NVIDIA researchers Yifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang, Huan Ling, Sanja Fidler, and Xuanchi Ren, reframes latent decoding as conditional pixel diffusion instead of reconstruction, so a latent from a VAE, SD3, DINOv2, or SigLIP is decoded directly into the final high-resolution image. Built on the 1.3B-parameter PixelDiT prior, PiD adds a lightweight ControlNet-style latent adapter and a sigma-aware gate that downweights noisy latents, enabling decoding from partially denoised LDM states and even early termination of the base sampler. The system is trained on 2.6M filtered high-resolution images from MultiAspect-4K-1M, rendered PDFs, and internal data, then distilled with DMD2 to only 4 denoising steps. In experiments, PiD decodes 512×512 latents into 2048×2048 outputs in under 1 second with 13 GB peak memory on an RTX 5090, and in 210 ms on a GB200 GPU, about 6× faster than cascaded diffusion super-resolution pipelines while improving visual quality. It also scales to 4K decoding, runs with less than 30 GB memory, and generalizes to semantic latents from DINOv2 and SigLIP, where it significantly improves no-reference quality metrics such as NIQE, MUSIQ, and VisualQuality-R1. The ablation study shows both the high-resolution text-to-image prior and sigma-aware gating are essential, and the 4-step student even surpasses multi-step teacher variants on perceptual metrics.
Original abstract
Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space, and a decoder maps the generated latents back to pixels. Yet the latent-to-pixel decoder is reconstruction-oriented, optimized to invert the encoder rather than synthesize more details, and becomes increasingly costly at megapixel scale. This drawback calls for a more expressive and efficient decoding paradigm. Motivated by recent progress in scalable pixel-space diffusion, we introduce PiD, a Pixel diffusion Decoder that reformulates latent decoding as conditional pixel diffusion, unifying decoding and upsampling into one generative module. By denoising directly in high-resolution pixel space, PiD synthesizes $4\times$ and even $8\times$ upscaled images with low latency. For latent conditioning, a lightweight sigma-aware adapter injects noise-corrupted latents into the pixel diffusion backbone, enabling PiD to decode partially denoised latents and terminate the latent diffusion process early. To further improve efficiency, we distill the model using DMD2, reducing inference to just 4 steps. PiD applies to both conventional VAE latents and semantic latents (e.g., SigLIP, DINOv2) used in recent RAE-based models. PiD decodes latents of $512 \times 512$ images into $2048 \times 2048$ pixels in under 1 second with 13 GB peak memory on a consumer RTX 5090, and as fast as 210 ms on a GB200 GPU, about $6\times$ faster than cascaded diffusion-based super-resolution pipelines with better visual fidelity.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.