Score-Control for Hallucination Reduction in Diffusion Models
AuthorsMahesh Bhosale, Naresh Kumar Devulapally, Abdul Wasi, Chau Pham, Vishnu Suresh Lokhande, David Doermann
This paper tackles hallucinations in diffusion image generators by controlling the model’s score dynamics, improving reliability without sacrificing quality.
Key results
Baseline hallucination rate on Hands-11K before VSM
Hallucination rate after applying VSM on Hands-11K
Baseline hallucination rate on the Cards benchmark
Hallucination rate after applying VSM on Cards
What the paper found
Score Control for Hallucination Reduction in Diffusion Models, from the University at Buffalo, argues that hallucinations in diffusion generators arise when the learned score field is overly smooth, which leaks nonzero probability mass into off-manifold regions. The paper formalizes this with a lower bound on model density outside data support that depends on the score Lipschitz constant L and score magnitude bound S, then introduces Variance-Guided Score Modulation, or VSM, an architecture-agnostic training objective that penalizes small score Jacobians using the variance-learning head from I-DDPM. A time-dependent schedule amplifies this penalty near the final denoising steps, where hallucinations are most likely. On synthetic and real image benchmarks, including 1D and 2D Gaussian mixtures, Hands-11K, MNIST, Shapes, Cards, ChessImages, and ImageNet-1K, VSM consistently reduces hallucination rates while preserving or improving fidelity and diversity. On the synthetic mixtures, it lowers score RMSE and hallucination rate; on Hands-11K, it cuts hallucinations from 23.33% to 5.15% for DDPM and from 29.50% to 21.15% for text-conditioned LDM; on Cards, it reduces hallucinations from 22.41% to 2.33%. The paper also introduces ChessImages, with an extreme semantic space of about 10^44 valid board states, to enable rule-checkable hallucination analysis and show that VSM increases valid novel generations rather than merely memorizing training boards.
Original abstract
Diffusion models have emerged as the backbone of modern generative AI, powering advances in vision, language, audio and other modalities. Despite their success, they suffer from hallucinations, implausible samples that lie outside the support of true data distribution, which degrade reliability and trust. In this work, we first empirically confirm previously proposed hypothesis that score smoothness causes hallucinations in Image Generation diffusion models and provide a density-based perspective. We further formalize this notion by linking the hallucinations probability mass to lipschitz constant of the learned score function. Motivated by this, we introduce a Variance-Guided Score Modulation (VSM) strategy that controls the score Jacobian, in turn reducing score smoothness and better approximating the ground truth score that decreases hallucinations. Empirical results on synthetic and real-world datasets demonstrate that our approach reduces hallucinations (up to ~25%) while maintaining high fidelity and diversity, providing a principled step toward more reliable diffusion-based image generation. We also propose two benchmark datasets with extreme semantic variation for systematic hallucination evaluation. Code and Datasets are publicly available at https://github.com/bhosalems/VSM.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.