Colored Noise Diffusion Sampling
AuthorsHadar Davidson, Noam Issachar, Sagie Benaim
Resources
This paper improves diffusion image generation by replacing uniform noise with frequency-aware colored noise during sampling, making the process more efficient and producing sharper images.
Key results
unguided ImageNet-256 baseline
unguided ImageNet-256 with CNS
unguided ImageNet-256 baseline
unguided ImageNet-256 with CNS
unguided ImageNet-256 baseline
unguided ImageNet-256 with CNS
What the paper found
Colored Noise Diffusion Sampling introduces CNS, a training-free inference-time sampler for diffusion and flow-matching models that replaces uniform white-noise injection with timestep- and frequency-dependent noise coloring. The paper argues that standard SDE solvers waste a finite energy budget because diffusion trajectories resolve low frequencies early and high frequencies late, so CNS uses a progression matrix γ(f,t) to route variance toward structurally unresolved bands while enforcing global variance conservation. Across ImageNet-256, CNS improves SiT-XL/2 FID from 8.26 to 6.27, JiT-B/16 from 32.39 to 26.69, and JiT-H/16 from 11.88 to 8.31, with CFG gains that remain consistent across solver orders including Euler-Maruyama, Heun, SRK2, and SRK2S. The method also transfers to Black Forest Labs’ FLUX.1-dev and FLUX.2-klein text-to-image pipelines, improving DrawBench human-preference and alignment metrics without retraining. Ablations show that both global energy mis-scaling and temporal shuffling sharply degrade quality, confirming that CNS’s advantage comes from where and when noise is injected, not just how much noise is used.
Original abstract
Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-frequency global structures early and high-frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing the finite energy budget. In this work, we establish a mathematical framework that reconsiders SDE inference as a targeted, frequency-decoupled energy transfer. Leveraging this framework, we introduce Colored Noise Sampling (CNS), a novel, training-free stochastic solver. Rather than injecting uniform white noise, CNS utilizes a dynamic, timestep- and frequency-dependent schedule that more efficiently allocates injected energy toward structurally unresolved frequency bands. By actively exploiting the model's inherent spectral bias, CNS systematically steers the generated distribution toward the true data manifold. Extensive experiments demonstrate that CNS significantly outperforms standard ODE and SDE baselines as a strictly plug-and-play, inference-time sampler substitution across diverse architectures (SiT, JiT, FLUX). Compared to standard sampling on ImageNet-256, CNS achieves substantial unguided FID reductions, improving from 8.26 to 6.27 on SiT-XL/2, 32.39 to 26.69 on JiT-B/16, and 11.88 to 8.31 on JiT-H/16, while yielding consistent relative FID improvements with Classifier-Free Guidance. Project page is available at https://hadardavidson.github.io/CNS/.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.