Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
AuthorsZiyue Lin, Jiahe Hou, Hongyu Xia, Xinrui Xie, Feifei Wang, Yuyin Zhou, Wei Wang, Jiawei Liu, Liangqiong Qu
This paper improves image-to-image translation by splitting diffusion into two stages so one part aligns domains with noise and the other learns the actual semantic change more efficiently.
Key results
DRDD achieves an average SSIM of 0.916 on the All-in-One-5 unified image restoration benchmark.
DRDD achieves an average LPIPS of 0.073 on the All-in-One-5 benchmark.
The paper evaluates DRDD on the CDD-11 benchmark, which contains 11 different degradation conditions.
The paper reports the best and most stable performance when Gaussian noise intensity is in the 0.8 to 1.3 range.
What the paper found
Decoupled Residual Denoising Diffusion Models, or DRDD, from researchers at the University of Hong Kong, the Chinese Academy of Sciences, and the University of California, Santa Cruz, reframes image-to-image translation by separating Gaussian noise injection from semantic residual removal. The paper’s key insight is that fixed Gaussian noise is not only a manifold-lifting trick but also a “domain harmonizer” that reduces feature divergence across tasks and domains; the authors formalize this with a KL-divergence result and then exploit it by splitting diffusion into a stochastic noise-diffusion stage and a deterministic residual-diffusion stage. In the reverse process, DRDD first removes task-specific residuals inside a noise-carrying domain, then denoises to the clean target, instead of entangling both operations in one coupled sampler as in RDDM or I2SB. This design improves data efficiency because the denoising module can be trained only on abundant unpaired target-domain images. Across All-in-One-5, CDD-11, and a new multi-domain denoising benchmark called MNMD, DRDD is consistently state of the art or competitive: on All-in-One-5 it raises average SSIM to 0.916 and cuts LPIPS to 0.073, outperforming strong diffusion and non-diffusion baselines such as DA-CLIP, DiffuIR, VLUNet, and DFPIR. It also remains robust under severe data pruning, and the authors show the framework extends to DDPM, DDIM, and SDE-based diffusion backbones such as IR-SDE. The best performance appears when noise intensity is around 0.8 to 1.3, supporting the paper’s central claim that controlled noise can be an asset, not just a nuisance.
Original abstract
We propose Decoupled Residual Denoising Diffusion models (DRDD) for unified and data-efficient image-to-image (I2I) translation. While diffusion models have advanced I2I translation in terms of quality and diversity, we uncover a previously under-explored property in diffusion models. Crucially, beyond its conventional role of manifold lifting (i.e., moving data off low-dimensional manifolds), injecting Gaussian noise facilitates domain harmonization by implicitly aligning feature distributions across domains, a property particularly advantageous for unified I2I translation. However, existing diffusion models prematurely erode this harmonization effect, as noise and residuals are simultaneously removed in a single coupled diffusion process. To address this, DRDD decouples the diffusion process into two sequential and independent diffusion stages: (1) a stochastic noise diffusion for domain harmonization and manifold lifting, and (2) a deterministic residual diffusion that learns the core semantic mapping entirely within the fixed-noise domain. This decoupling preserves harmonization and manifold lifting effects throughout the transformation, substantially simplifying the learning of unified mappings across diverse tasks and domains. Notably, the noise diffusion stage is trained exclusively on abundant, unpaired target-domain images, greatly improving data efficiency. Comprehensive theoretical and empirical analysis demonstrates that DRDD is broadly compatible with mainstream diffusion models and consistently delivers robust, unified I2I translation, even under limited paired data. Our code is available at https://github.com/HKU-HealthAI/DRDD.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.