NTH

Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution

AuthorsZixin Jessie Chen, Zhuo Chen, Archer Wang, Jeff Gore, William T. Freeman, Congyue Deng, Marin Soljačić

June 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a single diffusion model that can both generate images and perform super-resolution by starting the denoising process at different scales.

Key results

2.65
CIFAR-10 FID

Best unconditional sample quality reported for SKILD on CIFAR-10.

9.63
CIFAR-10 Inception Score

Unconditional CIFAR-10 generation quality for SKILD.

2x-8x
ImageNet SR factor range

Continuous super-resolution range supported by one trained checkpoint.

0.186
ImageNet-256 LPIPS

Best perceptual similarity score reported for 4x super-resolution.

0.612
ImageNet-256 CLIPIQA

Best CLIP-based perceptual quality score reported for 4x super-resolution.

59.226
ImageNet-256 MUSIQ

Best multiscale image quality score reported for 4x super-resolution.

What the paper found

This MIT-led paper introduces SKILD, a Scale-invariant K-Space Image Learning Diffusion model that makes image scale an explicit diffusion coordinate by attenuating high-frequency DCT modes before low-frequency ones while injecting spectrum-matched Gaussian noise, so one unconditional reverse process can do both generation and continuous super-resolution with no task-specific conditioning, classifier-free guidance, or per-scale retraining. On CIFAR-10, the model reaches FID 2.65 and Inception Score 9.63, making it the strongest among the frequency-informed diffusion methods compared in the paper. The same checkpoint then performs continuous ImageNet super-resolution from 2× to 8×, and at 4× on ImageNet-256 it achieves the best LPIPS at 0.186, CLIPIQA at 0.612, and MUSIQ at 59.226, while remaining competitive on PSNR and SSIM. The authors also validate the scale-invariant mechanism on a 128×128 critical 2D Ising benchmark, where reconstructions from a 32×32 effective-resolution start preserve the connected four-point correlation κ4 across patch sizes far better than SR3. The key novelty is that the forward process preserves the dataset’s empirical power spectrum, which the paper shows follows an approximate k^-2 law across CIFAR-10 and ImageNet, allowing super-resolution to be interpreted as starting the same diffusion trajectory from different timesteps.

Original abstract

Creating images from noise is image generation; reconstructing fine details from coarse inputs is super-resolution. Despite their practical differences, both can be understood as reversing information loss across scales. We introduce $\textbf{SKILD}$, a $\textbf{S}$cale-invariant $\textbf{K}$-Space $\textbf{I}$mage $\textbf{L}$earning $\textbf{D}$iffusion model that unifies generation and continuous super-resolution within a single unconditional framework. Both natural images and critical physical systems exhibit scale invariance, and we leverage it to design a forward process that attenuates image content from fine to coarse scales while injecting spectrum-matched Gaussian noise, making scale an explicit coordinate of the diffusion dynamics. The same trained reverse process performs generation and continuous super-resolution by varying only the starting timestep: $\textit{no task-specific architecture, no conditioning branch, no classifier-free guidance, no retraining per scale factor}$. Empirically, SKILD reaches FID $2.65$ and Inception Score $9.63$ on unconditional CIFAR-10, performs $2\times$--$8\times$ super-resolution on ImageNet from a single unconditional checkpoint while outperforming conditional models across perceptual metrics, and reconstructs critical Ising models whose connected four-point correlations closely track the ground truth.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis