NTH

Efficient and Training-Free Single-Image Diffusion Models

AuthorsHaojun Qiu, Kiriakos N. Kutulakos, David B. Lindell

June 16, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns single-image diffusion into a fast, training-free process by using closed-form patch statistics, enabling high-quality image generation from one reference image in seconds rather than hours.

Key results

0.21
SIFID

Proposed method with T = 40, η = 1 on unconditional generation

0.48
SinDDM SIFID

Baseline single-image diffusion model in the same unconditional benchmark

0.50
LPIPS Diversity

Highest diversity reported for the proposed method

8
VAE compression

Latent-space acceleration uses 8× spatial downsampling

834
Gigapixel time

End-to-end generation time in seconds for a 1 GP image

1000
Speedup

Acceleration over naive implementation at 16 MP

What the paper found

Efficient and Training-Free Single-Image Diffusion Models from the University of Toronto and the Vector Institute replaces hours of single-image diffusion training with a closed-form patch denoiser derived from the finite set of overlapping patches in one reference image. The core idea is to model noisy patch likelihoods directly, so the optimal denoiser becomes a weighted average over all reference patches, equivalent to a patch-level Gaussian mixture posterior mean and implementable as attention. This yields a training-free reverse diffusion sampler that works at one scale or in a coarse-to-fine pyramid, preserving global layout while recovering fine texture. On unconditional generation, the method matches or exceeds trained baselines such as SinDDM, SinFusion, and SinDiffusion on image quality and diversity, with SIFID improving to 0.21 from 0.48 for SinDDM in one setting and LPIPS diversity reaching 0.50. The paper also shows controllable symmetrization, retargeting, text-guided stylization with CLIP ViT-B/32, and structural analogy, while accelerating inference via fused attention, latent-space diffusion with an 8× VAE compression, and approximate nearest neighbors. These optimizations enable 1 GP generation in 834 s, 13.9 minutes for a 1 GP image, and, at 16 MP, more than 1000× speedup over a naive implementation.

Original abstract

We consider the problem of generating images whose internal structure -- defined by the distribution of patches across multiple scales -- matches that of a single reference image. Recent approaches address this problem by training a diffusion model on a single image. But even in this setting, training is computationally expensive and requires hours of optimization. Instead, we model the image using a dataset of its patches at different scales. As this dataset is finite and the dimensionality of its patches is small, the score function for a noisy patch can be computed tractably using an optimal, closed-form denoiser, eliminating the need for neural network training. We integrate this patch-based denoiser into an efficient, training-free image diffusion model, and we describe how our method connects to classical patch-based image restoration techniques. Our approach achieves state-of-the-art generation quality and diversity compared to trained single-image diffusion models, and we demonstrate applications, including unconditional image generation, text-guided stylization, image symmetrization, and retargeting. Further, we show that our approach is compatible with latent space diffusion, and we show multiple additional acceleration techniques to achieve megapixel single-image generation in one second, and gigapixel generation in minutes.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis