NTH

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

AuthorsSongchun Zhang, Sitong Guo, Xianghao Kong, Pengwei Liu, Yuwei Guo, Lvmin Zhang, Anyi Rao

August 7, 2026 2 min read
Watch on YouTube
The one-line take

FlexComposer makes it easier to place still or moving objects into videos while preserving realistic motion, appearance, and user-defined trajectories.

Key results

14B
Backbone size

Wan2.1-I2V-14B diffusion-transformer backbone.

12k
Synthetic training clips

Procedural simulation clips used for geometric bootstrap training.

54k
Real-world training clips

Real-world clips used for adaptation and photometric harmonization.

5k
Generative training clips

Generative clips used for open-domain refinement.

91.20%
Dynamic subject consistency

FlexComposer’s VBench subject-consistency score.

468.30
Dynamic compositing FVD

VBench dynamic-compositing FVD, compared with 535.71 for GenCompositor.

What the paper found

FlexComposer, from researchers at HKUST, Zhejiang University, CUHK, and Stanford University, reframes video compositing as trajectory-guided conditional generation, allowing either a static image or dynamic footage to be inserted into a target video while preserving identity and intrinsic motion. Its Unified Canonical Foreground Representation stabilizes and centers the asset, separating local dynamics from global displacement; Spatial-Aware Latent Injection then transports those features along user-defined 3D trajectories directly in VAE latent space, without ControlNet or other learnable adapters. A visibility gate handles occlusions, while static-image expansion with temporal noise encourages plausible motion and relighting augmentation teaches lighting and shadow adaptation. The system fine-tunes the Wan2.1-I2V-14B diffusion-transformer backbone with LoRA, trained on 12k synthetic clips, 54k real-world clips, and 5k generative clips through a synthetic-to-real curriculum on NVIDIA A100 GPUs. On dynamic compositing, FlexComposer reaches 91.20% subject consistency, 93.85% background consistency, and 98.71% motion smoothness on VBench, while reducing FVD to 468.30 versus 535.71 for GenCompositor. On DAVIS, its full-video V2V variant records an EPE of 2.15 and FVD of 82.6, improving trajectory adherence and temporal quality over Wan-Move. Compared with commercial systems such as Kling 1.5, the method prioritizes controllable placement and asset fidelity, though rapid motions can still produce blur and physical interactions remain diffusion-generated rather than explicitly simulated.

Original abstract

Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis