NTH

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

AuthorsJianzong Wu, Hao Lian, Jiongfan Yang, Dachao Hao, Ye Tian, Yunhai Tong, Jingyuan Zhu, Biaolong Chen, Qiaosong Qi, Aixi Zhang, Wanggui He, Mushui Liu, Jinlong Liu, Hao Jiang

June 6, 2026 2 min read
Watch on YouTube
The one-line take

LoomVideo is a faster, smaller video generation-and-editing model that uses a clever conditioning trick to avoid expensive token concatenation while still handling multimodal inputs.

Key results

5B
Model size

LoomVideo is described as a compact unified video generation and editing model built on a 5B-parameter architecture.

8B
Backbone VLM size

The model initialization uses Qwen3-VL-8B-Instruct as the multimodal language model backbone.

5.41×
Inference speedup

The zero-overhead Scale-and-Add conditioning achieves at least a 5.41× faster inference speed for video editing compared with similar concatenation-based models.

What the paper found

LoomVideo, developed by Peking University and Alibaba Group, is a 5B-parameter unified video foundation model for both generation and editing that targets interleaved multimodal inputs such as text, reference images, and source videos. Its key novelty is replacing the standard text encoder with Qwen3-VL-8B-Instruct and using a Deepstack injection mechanism that extracts hidden states from every MLLM layer and injects them into the corresponding Diffusion Transformer layers via cross-attention, improving semantic alignment without heavy adapters. For editing, LoomVideo avoids the token-concatenation strategy used by prior unified models like UniVideo and VINO; instead, it introduces a zero-overhead Scale-and-Add conditioning scheme that directly adds a timestep-scaled clean source latent to the target noisy latent, eliminating sequence-length blowup and enabling at least 5.41× faster inference. It also proposes Negative Temporal RoPE indices to cleanly distinguish multiple reference images from video frames. Trained in three stages on roughly O(10M) image-text pairs, O(10M) video-text pairs, O(10M) image-edit samples, O(3M) video-edit samples, plus reference-guided and multi-reference data, and further improved with DiffusionNFT reinforcement learning using PickScore, LoomVideo achieves state-of-the-art or highly competitive results on VBench, OpenVE-Bench, RefVIE-Bench, and IntelligentVBench, with especially strong gains in e-commerce and fashion scenarios on the new FashionVideoBench.

Original abstract

Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified frameworks predominantly rely on massive models (typically 13B parameters or more) and incorporate source video conditions for editing by concatenating sequence tokens. This concatenation inevitably doubles the sequence length, quadrupling the computational complexity of the self-attention mechanism and introducing prohibitive overhead. To address these bottlenecks, we present LoomVideo, a highly efficient 5B-parameter unified architecture for both video generation and editing. LoomVideo replaces the standard text encoder with a Multimodal Large Language Model (MLLM) and employs Deepstack injection mechanism to align multi-layer MLLM features with the Diffusion Transformer (DiT). Crucially, we introduce a zero-overhead Scale-and-Add conditioning approach for video editing. By scaling and directly adding the clean source video latent to the noised target latent, this elegant design eliminates the need for token concatenation, drastically reducing computational cost while maintaining robust capabilities for complex, non-rigid edits. Furthermore, a Negative Temporal RoPE strategy is seamlessly integrated to handle multiple reference images. Extensive experiments demonstrate that our compact 5B model achieves state-of-the-art or highly competitive performance across comprehensive benchmarks, exhibiting exceptional superiority in e-commerce and fashion generation scenarios. Benefiting from the zero-overhead conditioning mechanism, LoomVideo achieves at least a 5.41x acceleration in inference speed compared to models of similar capabilities, paving the way for highly practical and efficient video foundation models.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis