NTH

DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence

AuthorsXu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou

AffiliationsPeking University · Singapore University of Technology and Design

October 10, 2026 2 min read
Watch on YouTube
The one-line take

DC-SAE compresses images aggressively while preserving detail, helping diffusion models train faster and generate high-quality results.

Key results

32
Spatial compression

Compression ratio used for the ImageNet 512 × 512 results.

29.79
PSNR

ImageNet 512 × 512 reconstruction quality.

3.37
gFID

ImageNet 512 × 512 generation score.

4.41
Autoencoder throughput speedup

Speedup over DC-AE at 1024 × 1024 on an NVIDIA H200.

0.84
GenEval

Text-to-image score for the 1.6B-parameter DiT with Qwen3-1.7B.

86.007
DPG-Bench

Text-to-image score for the 1.6B-parameter DiT with Qwen3-1.7B.

What the paper found

DC-SAE tackles the tradeoff between compact image latents and fast diffusion training with a two-branch autoencoder: a frozen semantic encoder such as DINOv2 supplies structured, generation-friendly features, while a trainable pixel encoder preserves texture, color, and fine detail. The branches are jointly trained, and a decoder-side Spatial DeMerger expands tokens only for reconstruction, keeping the diffusion model’s latent grid compact. On ImageNet at 512 × 512, 32× spatial compression yields 29.79 PSNR and 3.37 gFID, compared with DC-AE’s 26.25 PSNR and 7.47 gFID; the paper reports improvements of 13.5% and 54.9%, respectively. At 1024 × 1024, DC-SAE delivers 4.41× the autoencoder throughput of DC-AE, measured on an NVIDIA H200. For text-to-image generation, a 1.6B-parameter DiT paired with a Qwen3-1.7B text encoder scores 0.84 on GenEval and 86.007 on DPG-Bench at 1024 × 1024. The results suggest that combining semantic structure with pixel-level detail can support both high compression and faster diffusion convergence.

Original abstract

High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a $1.6$B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at $1024\times1024$ resolution.

Read the original paper

More in Diffusion Models

Browse all 61 papers →
01Diffusion

ALoDLM: Adaptively Looped Diffusion Language Models

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao, Rajat Koner, Jiaye Wu, Linghan Xu, Xuanbai Chen, Xiang Xu, Zheng Zhang, Jakub Zablocki, Nishant Sankaran, Yifan Xing

ALoDLM makes diffusion language models smarter and faster by spending extra computation only on tokens that are difficult to predict.

Read analysis
02Diffusion

Learning to Read the Contextual Tokens in Diffusion Transformers

Omer Dahary, Etai Sella, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik

The paper teaches an LLM to read what image-generation models are thinking mid-generation, revealing hidden visual semantics and using them to improve diffusion outputs.

Read analysis