NTH

Score-Control for Hallucination Reduction in Diffusion Models

AuthorsMahesh Bhosale, Naresh Kumar Devulapally, Abdul Wasi, Chau Pham, Vishnu Suresh Lokhande, David Doermann

June 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper tackles hallucinations in diffusion image generators by controlling the model’s score dynamics, improving reliability without sacrificing quality.

Key results

23.33%
Hands-11K DDPM hallucination rate

Baseline hallucination rate on Hands-11K before VSM

5.15%
Hands-11K DDPM + VSM hallucination rate

Hallucination rate after applying VSM on Hands-11K

22.41%
Cards DDPM hallucination rate

Baseline hallucination rate on the Cards benchmark

2.33%
Cards DDPM + VSM hallucination rate

Hallucination rate after applying VSM on Cards

What the paper found

Score Control for Hallucination Reduction in Diffusion Models, from the University at Buffalo, argues that hallucinations in diffusion generators arise when the learned score field is overly smooth, which leaks nonzero probability mass into off-manifold regions. The paper formalizes this with a lower bound on model density outside data support that depends on the score Lipschitz constant L and score magnitude bound S, then introduces Variance-Guided Score Modulation, or VSM, an architecture-agnostic training objective that penalizes small score Jacobians using the variance-learning head from I-DDPM. A time-dependent schedule amplifies this penalty near the final denoising steps, where hallucinations are most likely. On synthetic and real image benchmarks, including 1D and 2D Gaussian mixtures, Hands-11K, MNIST, Shapes, Cards, ChessImages, and ImageNet-1K, VSM consistently reduces hallucination rates while preserving or improving fidelity and diversity. On the synthetic mixtures, it lowers score RMSE and hallucination rate; on Hands-11K, it cuts hallucinations from 23.33% to 5.15% for DDPM and from 29.50% to 21.15% for text-conditioned LDM; on Cards, it reduces hallucinations from 22.41% to 2.33%. The paper also introduces ChessImages, with an extreme semantic space of about 10^44 valid board states, to enable rule-checkable hallucination analysis and show that VSM increases valid novel generations rather than merely memorizing training boards.

Original abstract

Diffusion models have emerged as the backbone of modern generative AI, powering advances in vision, language, audio and other modalities. Despite their success, they suffer from hallucinations, implausible samples that lie outside the support of true data distribution, which degrade reliability and trust. In this work, we first empirically confirm previously proposed hypothesis that score smoothness causes hallucinations in Image Generation diffusion models and provide a density-based perspective. We further formalize this notion by linking the hallucinations probability mass to lipschitz constant of the learned score function. Motivated by this, we introduce a Variance-Guided Score Modulation (VSM) strategy that controls the score Jacobian, in turn reducing score smoothness and better approximating the ground truth score that decreases hallucinations. Empirical results on synthetic and real-world datasets demonstrate that our approach reduces hallucinations (up to ~25%) while maintaining high fidelity and diversity, providing a principled step toward more reliable diffusion-based image generation. We also propose two benchmark datasets with extreme semantic variation for systematic hallucination evaluation. Code and Datasets are publicly available at https://github.com/bhosalems/VSM.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis