Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference
AuthorsAgata Żywot, Iason Skylitsis, Thijmen Nijdam, Zoe Tzifa-Kratira, Derck Prinzhorn, Konrad Szewczyk, Aritra Bhowmik
Resources
This paper lets Stable Diffusion take both a text prompt and a reference image at inference time, so you can steer generation toward a desired style or visual concept without retraining.
Key results
Training data fraction used for the aligner
Aligner training completes in under this many hours on a single A100 GPU
Stable Diffusion v2.1 generation resolution
Sampling steps used in all experiments
Text alignment metric reported for the baseline and compared methods
Reference-image perceptual similarity metric reported for VCF
What the paper found
The paper introduces Visual Concept Fusion, a training-free inference-time method for injecting reference images into text-conditioned diffusion models such as Stable Diffusion v2 without retraining the generator or concept-specific fine-tuning. The core idea is to align CLIP image tokens to the CLIP text-token manifold with a lightweight 2-layer MLP trained on only a 10% subset of COCO Captions, then fuse aligned image and text tokens through concatenation, naive blending, or cross-attention; an optional Prompt-Noise Optimization module further refines the initial DDIM noise and conditioning tokens. Using Stable Diffusion v2.1 at 768×768 with 50 DDIM steps, the method preserves prompt semantics while transferring style, composition, and color palette from the reference image. Quantitatively, the main trade-off is between CLIP text alignment and LPIPS reference similarity: text-only SDv2 scores 0.29 CLIP and 0.78 LPIPS, naive fusion scores 0.28 and 0.77, and VCF reaches 0.27 and 0.76, indicating stronger perceptual adherence to the reference at a modest cost in prompt literalism. The aligner is lightweight, trains in under two hours on a single A100 GPU, and ablations show that combining InfoNCE with cross-attention reconstruction is necessary for balancing global semantic alignment and local visual detail; Prompt-Noise Optimization further reduces noise and improves fidelity to reference-specific features.
Original abstract
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally expensive fine-tuning or rely on style transfer techniques that risk semantic misalignment with textual prompts. We introduce Visual Concept Fusion (VCF), the first method offering dual conditioning on both an image and text prompt at inference time without any concept-specific training. VCF enables visual concept injection into Stable Diffusion by aligning CLIP image features with the text embedding space. VCF consists of three components: (1) a lightweight aligner that maps image tokens to the text embedding manifold using InfoNCE and cross-attention reconstruction losses, (2) a fusion strategy that preserves both textual and visual semantics, and (3) an optional Prompt-Noise Optimization (PNO) module for test-time refinement. Our experiments demonstrate that VCF successfully transfers visual attributes including style, composition, and color palette from reference images while maintaining prompt adherence. Quantitative results show a trade-off between text alignment (CLIP score) and visual correspondence (LPIPS), with VCF outperforming baselines in reference fidelity.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.