Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution
AuthorsZixin Jessie Chen, Zhuo Chen, Archer Wang, Jeff Gore, William T. Freeman, Congyue Deng, Marin Soljačić
Resources
This paper introduces a single diffusion model that can both generate images and perform super-resolution by starting the denoising process at different scales.
Key results
Best unconditional sample quality reported for SKILD on CIFAR-10.
Unconditional CIFAR-10 generation quality for SKILD.
Continuous super-resolution range supported by one trained checkpoint.
Best perceptual similarity score reported for 4x super-resolution.
Best CLIP-based perceptual quality score reported for 4x super-resolution.
Best multiscale image quality score reported for 4x super-resolution.
What the paper found
This MIT-led paper introduces SKILD, a Scale-invariant K-Space Image Learning Diffusion model that makes image scale an explicit diffusion coordinate by attenuating high-frequency DCT modes before low-frequency ones while injecting spectrum-matched Gaussian noise, so one unconditional reverse process can do both generation and continuous super-resolution with no task-specific conditioning, classifier-free guidance, or per-scale retraining. On CIFAR-10, the model reaches FID 2.65 and Inception Score 9.63, making it the strongest among the frequency-informed diffusion methods compared in the paper. The same checkpoint then performs continuous ImageNet super-resolution from 2× to 8×, and at 4× on ImageNet-256 it achieves the best LPIPS at 0.186, CLIPIQA at 0.612, and MUSIQ at 59.226, while remaining competitive on PSNR and SSIM. The authors also validate the scale-invariant mechanism on a 128×128 critical 2D Ising benchmark, where reconstructions from a 32×32 effective-resolution start preserve the connected four-point correlation κ4 across patch sizes far better than SR3. The key novelty is that the forward process preserves the dataset’s empirical power spectrum, which the paper shows follows an approximate k^-2 law across CIFAR-10 and ImageNet, allowing super-resolution to be interpreted as starting the same diffusion trajectory from different timesteps.
Original abstract
Creating images from noise is image generation; reconstructing fine details from coarse inputs is super-resolution. Despite their practical differences, both can be understood as reversing information loss across scales. We introduce $\textbf{SKILD}$, a $\textbf{S}$cale-invariant $\textbf{K}$-Space $\textbf{I}$mage $\textbf{L}$earning $\textbf{D}$iffusion model that unifies generation and continuous super-resolution within a single unconditional framework. Both natural images and critical physical systems exhibit scale invariance, and we leverage it to design a forward process that attenuates image content from fine to coarse scales while injecting spectrum-matched Gaussian noise, making scale an explicit coordinate of the diffusion dynamics. The same trained reverse process performs generation and continuous super-resolution by varying only the starting timestep: $\textit{no task-specific architecture, no conditioning branch, no classifier-free guidance, no retraining per scale factor}$. Empirically, SKILD reaches FID $2.65$ and Inception Score $9.63$ on unconditional CIFAR-10, performs $2\times$--$8\times$ super-resolution on ImageNet from a single unconditional checkpoint while outperforming conditional models across perceptual metrics, and reconstructs critical Ising models whose connected four-point correlations closely track the ground truth.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.