NTH

Pixel-Space Diffusion via Observation Operators

AuthorsShaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang

September 2, 2026 2 min read
Watch on YouTube
The one-line take

This work makes pixel-space diffusion easier to train by teaching models to recover images from coarse structures to fine details in the same order humans and vision systems naturally perceive them.

Key results

1.52
ImageNet-256 gFID

Best reported generation quality after extended training.

260
Extended training

Epochs used to reach gFID 1.52 at 256×256 resolution.

1.95
Early ImageNet-256 gFID

Score achieved after only 80 training epochs.

1.58
ImageNet-512 gFID

Result at 512×512 resolution.

12.7%
Improvement over prior pixel-space result

gFID improvement at 512×512 resolution.

798M
Model parameters

Parameter count of the proposed model.

What the paper found

Pixel-space diffusion avoids the compression artifacts of latent models such as Stable Diffusion, but it is harder to optimize because conventional flow matching supervises the entire clean image at every noise level. This paper identifies that conflict as a scale–time mismatch: coarse structure is predictable under heavy noise, while fine detail emerges later. Observation Operator Diffusion replaces the fixed target with a time-indexed trajectory of Gaussian–Lanczos observations, smoothly progressing from low-pass structure to the full image, and adds the derivative of that trajectory to the velocity target. Its GL-CoDA decoder extends the same coarse-to-fine principle through network depth by injecting scale-specific structural responses into a Patch-DiT backbone, while retaining REPA representation alignment and direct v prediction. On class-conditional ImageNet-1K, the 798M-parameter model reaches gFID 1.95 after 80 epochs and 1.52 after 260 epochs at 256×256 resolution, using 60 fewer epochs than the prior best pixel-space result. At 512×512, it achieves gFID 1.58, a 12.7% improvement over the previous pixel-space benchmark, with recall rising to 0.69. Ablations show the combined Gaussian–Lanczos kernel outperforming identity, Gaussian-only, and Lanczos-only supervision, supporting the claim that scale-aligned targets improve gradient signal and convergence. The work is particularly relevant to industrial image-generation efforts, including those associated with JD.com, but it does not depend on proprietary systems such as OpenAI’s Sora or Meta’s Llama.

Original abstract

Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while still relying on a fixed clean-image target throughout denoising. Through empirical analysis, we identify a scale-time mismatch: image structures become predictable from coarse to fine as noise decreases, whereas existing models are forced to predict the full image even under high noise, resulting in low-SNR gradients that hinder optimization. To resolve this mismatch, we propose Observation Operator Diffusion, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures. Specifically, we replace fixed full-image supervision along the standard flow path with a time-indexed observation trajectory that evolves from coarse structures to the full image during denoising. This trajectory is instantiated with a family of Gaussian-Lanczos operators at varying observation scales, yielding a path-consistent training objective. We further introduce GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine feature refinement. Extensive experiments show that the proposed approach converges substantially faster while consistently improving generation quality, achieving an FID of 1.52 on ImageNet-256.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis