Spectral Progressive Diffusion for Efficient Image and Video Generation
AuthorsHoward Xiao, Brian Chao, Lior Yariv, Gordon Wetzstein
Resources
This paper speeds up diffusion-based image and video generation by growing resolution gradually in the frequency domain, so models spend computation where it actually matters.
Key results
Training-free acceleration on FLUX.1-dev at 1024² reaches this wall-clock speedup in the 7-step, S = 3 setting while preserving competitive image-quality and text-alignment metrics.
The largest FLUX.1-dev setting reduces total FLOPs by 7.36× over the full-resolution baseline.
Training-free latent-space video generation on WAN 2.1 at 720P achieves this wall-clock speedup in the S = 3 setting.
The fine-tuning recipe with LoRA rank 32 improves Z-Image latent image generation to this speedup while maintaining matched or better prompt alignment than the 50-step baseline.
The fine-tuning experiments use LoRA with rank 32 to bridge the training-inference gap for progressive-resolution generation.
The measured power spectrum fit for FLUX.1-dev latents is used to derive the δ-optimal resolution schedule from the model’s spectral decay.
What the paper found
Spectral Progressive Diffusion introduces a training-free and fine-tunable acceleration framework for pretrained diffusion and flow-matching generators by exploiting spectral autoregression: low-frequency structure emerges earlier than high-frequency detail, so the method keeps early denoising in a smaller token grid and expands resolution only when the model’s power spectrum predicts that finer frequencies become signal-dominated. Its core mechanism, spectral noise expansion, uses an orthonormal transform such as DCT to carry partially denoised low-frequency coefficients into a larger grid while injecting newly available high-frequency coefficients with correctly scaled Gaussian noise, then re-aligns the timestep analytically. The paper derives a δ-optimal transition schedule from the measured power spectrum, with fitted exponents β = 1.9155 for FLUX.1-dev latents, 2.4227 for WAN 2.1 video latents, and 2.4493 for PixelGen pixels, so the schedule depends on a single error-tolerance hyperparameter rather than brittle manual heuristics. On FLUX.1-dev at 1024², the method reaches 7.09× wall-clock speedup and 7.36× FLOPs reduction while preserving competitive ImageReward, CLIP-IQA, NIQE, T2I-CompBench, and GenEval scores; on WAN 2.1 video generation at 720p, it delivers 2.54× speedup. The fine-tuning recipe, implemented with LoRA rank 32, narrows the training-inference gap further and on Z-Image improves the 50-step baseline to 7.81× speedup at matched or better prompt alignment, while also enabling frequency-based image editing that outperforms SDEdit in geometric consistency and stylization fidelity.
Original abstract
Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later timesteps. This structure offers a natural opportunity for efficient generation, as high-resolution computation on noise-dominated frequencies is largely redundant. We propose Spectral Progressive Diffusion, a general framework that progressively grows resolution along the denoising trajectory of pretrained diffusion models. To this end, we develop a spectral noise expansion mechanism and derive an optimal resolution schedule from the model's power spectrum. Our framework supports training-free acceleration and a novel fine-tuning recipe that further improves efficiency and quality. We demonstrate significant speedups on state-of-the-art pretrained image and video generation models while preserving visual quality.
Read the original paper