NTH

Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference

AuthorsAgata Żywot, Iason Skylitsis, Thijmen Nijdam, Zoe Tzifa-Kratira, Derck Prinzhorn, Konrad Szewczyk, Aritra Bhowmik

June 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper lets Stable Diffusion take both a text prompt and a reference image at inference time, so you can steer generation toward a desired style or visual concept without retraining.

Key results

10%
COCO Captions subset

Training data fraction used for the aligner

2
Training time

Aligner training completes in under this many hours on a single A100 GPU

768x768
Image resolution

Stable Diffusion v2.1 generation resolution

50
DDIM steps

Sampling steps used in all experiments

0.29
CLIP score (SDv2 / naive / VCF)

Text alignment metric reported for the baseline and compared methods

0.76
LPIPS (SDv2 / naive / VCF)

Reference-image perceptual similarity metric reported for VCF

What the paper found

The paper introduces Visual Concept Fusion, a training-free inference-time method for injecting reference images into text-conditioned diffusion models such as Stable Diffusion v2 without retraining the generator or concept-specific fine-tuning. The core idea is to align CLIP image tokens to the CLIP text-token manifold with a lightweight 2-layer MLP trained on only a 10% subset of COCO Captions, then fuse aligned image and text tokens through concatenation, naive blending, or cross-attention; an optional Prompt-Noise Optimization module further refines the initial DDIM noise and conditioning tokens. Using Stable Diffusion v2.1 at 768×768 with 50 DDIM steps, the method preserves prompt semantics while transferring style, composition, and color palette from the reference image. Quantitatively, the main trade-off is between CLIP text alignment and LPIPS reference similarity: text-only SDv2 scores 0.29 CLIP and 0.78 LPIPS, naive fusion scores 0.28 and 0.77, and VCF reaches 0.27 and 0.76, indicating stronger perceptual adherence to the reference at a modest cost in prompt literalism. The aligner is lightweight, trains in under two hours on a single A100 GPU, and ablations show that combining InfoNCE with cross-attention reconstruction is necessary for balancing global semantic alignment and local visual detail; Prompt-Noise Optimization further reduces noise and improves fidelity to reference-specific features.

Original abstract

Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally expensive fine-tuning or rely on style transfer techniques that risk semantic misalignment with textual prompts. We introduce Visual Concept Fusion (VCF), the first method offering dual conditioning on both an image and text prompt at inference time without any concept-specific training. VCF enables visual concept injection into Stable Diffusion by aligning CLIP image features with the text embedding space. VCF consists of three components: (1) a lightweight aligner that maps image tokens to the text embedding manifold using InfoNCE and cross-attention reconstruction losses, (2) a fusion strategy that preserves both textual and visual semantics, and (3) an optional Prompt-Noise Optimization (PNO) module for test-time refinement. Our experiments demonstrate that VCF successfully transfers visual attributes including style, composition, and color palette from reference images while maintaining prompt adherence. Quantitative results show a trade-off between text alignment (CLIP score) and visual correspondence (LPIPS), with VCF outperforming baselines in reference fidelity.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis