Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
AuthorsWeichen Fan, Haiwen Diao, Penghao Wu, Ziwei Liu
Resources
This paper boosts pixel-space diffusion by filtering out high-frequency noise early in the denoising process, helping the model focus on the signal where it matters most.
Key results
JiT-700M/32 at 60 epochs without Spectral Forcing
JiT-700M/32 at 60 epochs with Spectral Forcing
JiT-700M/32 at 60 epochs without Spectral Forcing
JiT-700M/32 at 60 epochs with Spectral Forcing
JiT-700M/32 on ImageNet-256, showing earlier convergence
What the paper found
Show the Signal, Hide the Noise introduces Spectral Forcing, a parameter-free input adapter for pixel-space rectified-flow diffusion that applies a time-conditional 2D-DCT low-pass mask before the patch embedder, expanding its cutoff toward the identity at the data endpoint. The paper’s core claim is that pixel diffusion wastes capacity outside a frequency-time wedge where the denoiser is only learning deterministic baselines rather than modeling the data distribution. On ImageNet-256, using JiT-700M/32 with 64 tokens, Spectral Forcing reduces FID from 24.19 to 20.68 and raises Inception Score from 83.28 to 93.96 in a 60-epoch apples-to-apples comparison, with gains that persist across checkpoints and reach 15.15 FID at 120 epochs, matching a previously published roughly 145-epoch reference sooner. The method is regime-dependent: it helps most when patch tokenization is coarse and high-frequency content is mostly noise, remains competitive at 256 tokens, and can hurt when high frequencies carry essential signal. The same unchanged operator transfers to SenseNova-U1, a unified text-to-image model from the NTU-linked authors, improving DPG-Bench and GenEval, which suggests the spectral prior is useful beyond class-conditional generation.
Original abstract
Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffusion and natural-image power-law spectra, the per-band data-to-noise contour $k^{*}(t) = (1-t)^{-2/α}$ separates a signal-bearing low-frequency region from a noise-dominated high-frequency region at each time $t$. We show that this implicit coarse-to-fine structure is not merely descriptive: it induces a capacity-allocation problem. A standard pixel-space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency-time regions where the optimal prediction collapses to deterministic baselines rather than data-distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter-free, time-conditional 2D-DCT low-pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high-frequency content is predominantly noise rather than essential signal. On ImageNet-256 with JiT-700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova-U1, a unified text-to-image model, where it improves DPG-Bench and GenEval, showing that the input-side spectral prior transfers beyond class-conditional generation. These results suggest a route to capacity-efficient pixel-space diffusion by showing the signal and hiding the noise.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.