Pixel-Space Diffusion via Observation Operators
AuthorsShaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang
Resources
This work makes pixel-space diffusion easier to train by teaching models to recover images from coarse structures to fine details in the same order humans and vision systems naturally perceive them.
Key results
Best reported generation quality after extended training.
Epochs used to reach gFID 1.52 at 256×256 resolution.
Score achieved after only 80 training epochs.
Result at 512×512 resolution.
gFID improvement at 512×512 resolution.
Parameter count of the proposed model.
What the paper found
Pixel-space diffusion avoids the compression artifacts of latent models such as Stable Diffusion, but it is harder to optimize because conventional flow matching supervises the entire clean image at every noise level. This paper identifies that conflict as a scale–time mismatch: coarse structure is predictable under heavy noise, while fine detail emerges later. Observation Operator Diffusion replaces the fixed target with a time-indexed trajectory of Gaussian–Lanczos observations, smoothly progressing from low-pass structure to the full image, and adds the derivative of that trajectory to the velocity target. Its GL-CoDA decoder extends the same coarse-to-fine principle through network depth by injecting scale-specific structural responses into a Patch-DiT backbone, while retaining REPA representation alignment and direct v prediction. On class-conditional ImageNet-1K, the 798M-parameter model reaches gFID 1.95 after 80 epochs and 1.52 after 260 epochs at 256×256 resolution, using 60 fewer epochs than the prior best pixel-space result. At 512×512, it achieves gFID 1.58, a 12.7% improvement over the previous pixel-space benchmark, with recall rising to 0.69. Ablations show the combined Gaussian–Lanczos kernel outperforming identity, Gaussian-only, and Lanczos-only supervision, supporting the claim that scale-aligned targets improve gradient signal and convergence. The work is particularly relevant to industrial image-generation efforts, including those associated with JD.com, but it does not depend on proprietary systems such as OpenAI’s Sora or Meta’s Llama.
Original abstract
Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while still relying on a fixed clean-image target throughout denoising. Through empirical analysis, we identify a scale-time mismatch: image structures become predictable from coarse to fine as noise decreases, whereas existing models are forced to predict the full image even under high noise, resulting in low-SNR gradients that hinder optimization. To resolve this mismatch, we propose Observation Operator Diffusion, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures. Specifically, we replace fixed full-image supervision along the standard flow path with a time-indexed observation trajectory that evolves from coarse structures to the full image during denoising. This trajectory is instantiated with a family of Gaussian-Lanczos operators at varying observation scales, yielding a path-consistent training objective. We further introduce GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine feature refinement. Extensive experiments show that the proposed approach converges substantially faster while consistently improving generation quality, achieving an FID of 1.52 on ImageNet-256.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.