Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting
AuthorsGeorgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou
This work makes it easier to combine multiple LoRA styles and concepts in diffusion models by weighting each adapter based on how important its trigger words are in the prompt.
Key results
Average image-fidelity score on ComposLoRA
Average image-fidelity score on ComposLoRA
Average identity-preservation score on ComposLoRA
Average text-image alignment score on ComposLoRA
Final denoising steps reserved for character LoRAs in W-Switch
Trial share preferring W-Switch in the human study
What the paper found
This Kingston University London paper tackles a practical failure mode in Stable Diffusion LoRA personalization: when multiple concept-specific LoRA adapters are composed, naive weight or output averaging causes concept interference and identity loss. The authors propose two training-free decoding-time methods, W-Switch and W-Composite, that assign each LoRA a prompt-aware importance weight from the target prompt’s text-encoder embeddings, using either Prompt Ablation Weighting or Prompt Trigger Weighting; W-Composite blends all LoRA noise predictions with fixed normalized weights, while W-Switch allocates diffusion timesteps proportionally and preserves the character LoRA during the final 5 denoising steps to protect facial detail. On the ComposLoRA benchmark with Stable Diffusion v1.5 and the Realistic Vision V5.1 checkpoint, W-Switch achieves the strongest overall results, reaching 75.14 ICLIP, 50.74 IDINO, 53.06 IArcFace, and 36.54 TCLIP average scores, outperforming LoRA-Switch, LoRA-Composite, and CMLoRA. The paper also introduces a more faithful image-based evaluation protocol that uses SAM3 or FAN cropping plus CLIP, DINOv2, and ArcFace similarity against real reference images, avoiding the centroid bias of text-based or global-embedding metrics. In a MiniCPM-V assessment, W-Switch scores 8.641 average across element integration, spatial consistency, semantic accuracy, and aesthetic appeal, and a 16-participant user study preferred it in 47.32% of trials, with statistically significant gains over all baselines.
Original abstract
Low-Rank Adaptation (LoRA) successfully enables personalization in text-to-image generation by adapting pre-trained diffusion models to specific visual concepts and styles. However, extending such models to multi-concept customization remains challenging. Naively combining multiple LoRA weights or their outputs often leads to interference among concepts, resulting in degraded visual quality and reduced fidelity to the reference images of individual concepts. This paper proposes a simple yet effective approach for multi-concept customization by optimally combining the outputs of multiple LoRA modules. We leverage the relative importance of each concept during generation, as inferred from its corresponding prompt tokens and introduce two methods, W-Switch and W-Composite, that employ a prompt-aware importance weighting strategy in which each LoRA is weighted according to the semantic influence of its trigger words in the target prompt. In addition, we extend existing quantitative evaluation metrics by proposing a new image-based similarity evaluation framework that assesses image fidelity and identity preservation through comparisons between real-world reference images and automatically segmented concept regions from generated images. We evaluate our approach on the ComposLoRA testbed and demonstrate consistent improvements over existing state-of-the-art methods in terms of visual quality, identity preservation and compositionality. Qualitative evaluations, including a Large Language Model (LLM) based assessment and a user study, further validate the effectiveness of the proposed methods and align with the newly introduced quantitative image-based metrics. Our code is available at https://github.com/GeorgeTsoumplekas/Prompt-Aware-Multi-LoRA-Composition.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.