NTH

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

AuthorsYifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang, Huan Ling, Sanja Fidler, Xuanchi Ren

June 24, 2026 2 min read
Watch on YouTube
The one-line take

PiD replaces traditional latent decoders with a fast pixel-diffusion decoder that can turn compressed latents into high-resolution images much more efficiently and with better visual fidelity.

Key results

2.6M
training images

High-quality filtered dataset used to train PiD

1.3B
base prior parameters

PixelDiT checkpoint used as the high-resolution pixel diffusion prior

4
student inference steps

DMD2-distilled decoder step count

13
RTX 5090 peak memory

Peak GPU memory in GB for 2048×2048 decoding

210
GB200 latency

End-to-end decoding time in ms for 2048×2048 outputs

6
speedup over cascaded SR

Approximate speedup over diffusion-based super-resolution pipelines

What the paper found

PiD, from NVIDIA researchers Yifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang, Huan Ling, Sanja Fidler, and Xuanchi Ren, reframes latent decoding as conditional pixel diffusion instead of reconstruction, so a latent from a VAE, SD3, DINOv2, or SigLIP is decoded directly into the final high-resolution image. Built on the 1.3B-parameter PixelDiT prior, PiD adds a lightweight ControlNet-style latent adapter and a sigma-aware gate that downweights noisy latents, enabling decoding from partially denoised LDM states and even early termination of the base sampler. The system is trained on 2.6M filtered high-resolution images from MultiAspect-4K-1M, rendered PDFs, and internal data, then distilled with DMD2 to only 4 denoising steps. In experiments, PiD decodes 512×512 latents into 2048×2048 outputs in under 1 second with 13 GB peak memory on an RTX 5090, and in 210 ms on a GB200 GPU, about 6× faster than cascaded diffusion super-resolution pipelines while improving visual quality. It also scales to 4K decoding, runs with less than 30 GB memory, and generalizes to semantic latents from DINOv2 and SigLIP, where it significantly improves no-reference quality metrics such as NIQE, MUSIQ, and VisualQuality-R1. The ablation study shows both the high-resolution text-to-image prior and sigma-aware gating are essential, and the 4-step student even surpasses multi-step teacher variants on perceptual metrics.

Original abstract

Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space, and a decoder maps the generated latents back to pixels. Yet the latent-to-pixel decoder is reconstruction-oriented, optimized to invert the encoder rather than synthesize more details, and becomes increasingly costly at megapixel scale. This drawback calls for a more expressive and efficient decoding paradigm. Motivated by recent progress in scalable pixel-space diffusion, we introduce PiD, a Pixel diffusion Decoder that reformulates latent decoding as conditional pixel diffusion, unifying decoding and upsampling into one generative module. By denoising directly in high-resolution pixel space, PiD synthesizes $4\times$ and even $8\times$ upscaled images with low latency. For latent conditioning, a lightweight sigma-aware adapter injects noise-corrupted latents into the pixel diffusion backbone, enabling PiD to decode partially denoised latents and terminate the latent diffusion process early. To further improve efficiency, we distill the model using DMD2, reducing inference to just 4 steps. PiD applies to both conventional VAE latents and semantic latents (e.g., SigLIP, DINOv2) used in recent RAE-based models. PiD decodes latents of $512 \times 512$ images into $2048 \times 2048$ pixels in under 1 second with 13 GB peak memory on a consumer RTX 5090, and as fast as 210 ms on a GB200 GPU, about $6\times$ faster than cascaded diffusion-based super-resolution pipelines with better visual fidelity.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis