NTH

Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion

AuthorsWeichen Fan, Haiwen Diao, Penghao Wu, Ziwei Liu

June 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper boosts pixel-space diffusion by filtering out high-frequency noise early in the denoising process, helping the model focus on the signal where it matters most.

Key results

24.19
ImageNet-256 FID baseline

JiT-700M/32 at 60 epochs without Spectral Forcing

20.68
ImageNet-256 FID with SF

JiT-700M/32 at 60 epochs with Spectral Forcing

83.28
ImageNet-256 IS baseline

JiT-700M/32 at 60 epochs without Spectral Forcing

93.96
ImageNet-256 IS with SF

JiT-700M/32 at 60 epochs with Spectral Forcing

15.15
120-epoch FID with SF

JiT-700M/32 on ImageNet-256, showing earlier convergence

What the paper found

Show the Signal, Hide the Noise introduces Spectral Forcing, a parameter-free input adapter for pixel-space rectified-flow diffusion that applies a time-conditional 2D-DCT low-pass mask before the patch embedder, expanding its cutoff toward the identity at the data endpoint. The paper’s core claim is that pixel diffusion wastes capacity outside a frequency-time wedge where the denoiser is only learning deterministic baselines rather than modeling the data distribution. On ImageNet-256, using JiT-700M/32 with 64 tokens, Spectral Forcing reduces FID from 24.19 to 20.68 and raises Inception Score from 83.28 to 93.96 in a 60-epoch apples-to-apples comparison, with gains that persist across checkpoints and reach 15.15 FID at 120 epochs, matching a previously published roughly 145-epoch reference sooner. The method is regime-dependent: it helps most when patch tokenization is coarse and high-frequency content is mostly noise, remains competitive at 256 tokens, and can hurt when high frequencies carry essential signal. The same unchanged operator transfers to SenseNova-U1, a unified text-to-image model from the NTU-linked authors, improving DPG-Bench and GenEval, which suggests the spectral prior is useful beyond class-conditional generation.

Original abstract

Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffusion and natural-image power-law spectra, the per-band data-to-noise contour $k^{*}(t) = (1-t)^{-2/α}$ separates a signal-bearing low-frequency region from a noise-dominated high-frequency region at each time $t$. We show that this implicit coarse-to-fine structure is not merely descriptive: it induces a capacity-allocation problem. A standard pixel-space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency-time regions where the optimal prediction collapses to deterministic baselines rather than data-distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter-free, time-conditional 2D-DCT low-pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high-frequency content is predominantly noise rather than essential signal. On ImageNet-256 with JiT-700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova-U1, a unified text-to-image model, where it improves DPG-Bench and GenEval, showing that the input-side spectral prior transfers beyond class-conditional generation. These results suggest a route to capacity-efficient pixel-space diffusion by showing the signal and hiding the noise.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis