CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation
AuthorsFangtai Wu, Hailong Guo, Shijie Huang, Jiayi Song, Yubo Huang, Mushui Liu, Zhao Wang, Yunlong Yu, Jiaming Liu, Ruihua Huang
Resources
CollectionLoRA condenses dozens of specialized image-editing LoRAs into a single efficient model, making diffusion-based customization easier to deploy without losing visual quality.
Key results
number of effect LoRAs consolidated into one student
unlabeled general-domain images used for regularization
few-step generation setting used by the unified model
EffectBench score for the proposed method
bad case rate on EffectBench
extended scaling experiment for CollectionLoRA
What the paper found
CollectionLoRA, from Zhejiang University and Alibaba’s Qwen Applications Business Group, reframes customized diffusion editing as a multi-teacher on-policy distillation problem and compresses up to 50 effect-specific LoRAs plus few-step generation into one LoRA, eliminating routing latency, storage growth, and parameter conflicts that cause concept bleeding and style drift. Built on Qwen-Image-Edit-2509 with Distribution Matching Distillation, it adds Probabilistic Dual-Stream Routing to mix effect data with 20K general-domain images, Asymmetric Orthogonal Prompting to separate teacher and student concepts using VLM-rewritten prompts and orthogonal trigger words, and a Coarse-to-Fine Distillation Objective that combines flow-matching trajectory anchoring with Target Simulation to prevent distribution collapse. On EffectBench, its 8-step unified model reaches CLIP 0.727, DreamSim 0.425, EditReward 1.052, VSA 4.380, and BCR 0.087, outperforming the 50-in-1 flow-matching baseline and reducing deployment overhead to 0s/q routing latency with 100% routing accuracy for 10–50 LoRAs. The method scales to 180 effects, where storage drops to 2.2G×3 and routing accuracy remains 82%, and it also shows zero-shot composition: two learned effects can be chained at inference without additional training.
Original abstract
Customized image editing aims to equip pre-trained diffusion models with specific visual effects using limited paired data, typically via Low-Rank Adaptation (LoRA). As the number of desired effects grows, storing and dynamically loading numerous these effect LoRAs significantly increases deployment overhead. Furthermore, current pipelines typically cascade these effect LoRAs with acceleration modules for fast generation, which triggers severe parameter interference and results in concept bleeding and style degradation. We propose CollectionLoRA, a multi-teacher on-policy distillation framework capable of distilling the concepts of up to 50 different effect LoRAs along with few-step generation capabilities into a single LoRA. This fundamentally resolves the feature interference issue and significantly reduces deployment costs. Specifically, the method introduces (i) a Probabilistic Dual-Stream Routing mechanism that enables the model to randomly switch between data sources during training, effectively enhancing its generalization in unseen scenarios; (ii) an Asymmetric Orthogonal Prompting strategy to achieve concept isolation within the prompt space; (iii) a Coarse-to-Fine Distillation Objective to mitigate the distribution gap between the teacher and student models. Extensive evaluations show that CollectionLoRA distills all customized effects and few-step generation into a single LoRA, reducing deployment overhead while achieving concept fidelity comparable to or better than independently trained teacher models.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.