OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
AuthorsJiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen Liu, Pengfei Wan
Resources
OmniDirector makes video generation follow complex multi-shot camera moves by turning camera paths into visual motion grids and training a large controllable diffusion system on them.
Key results
internet videos used to build the camera grid–video dataset
portion of training samples using camera-grid reconstruction
curated validation samples
OmniDirector camera accuracy
OmniDirector T-Pre on the evaluation set
reference-video leakage at frame level
What the paper found
OmniDirector, developed by the Kuaishou Technology Kling team with Tsinghua University and Peking University collaborators, tackles general multi-shot camera cloning by replacing brittle camera-parameter injection and scarce cross-paired data with a visual camera grid: reference videos are decomposed into camera poses, rendered as motion in an empty 3D scene, and fed into a multimodal diffusion transformer as a spatiotemporal conditioning stream. The method is trained on a million-scale camera grid–video dataset built from 1.8M internet videos, then fine-tuned with a self-reconstruction objective on 30% of samples to force the model to learn geometry rather than appearance leakage. At inference, a hierarchical Prompt Expansion agent uses Qwen3-VL to generate inter-shot and intra-shot camera descriptions and fuse them with the user prompt and reference image, while adaptive classifier-free guidance separates camera structure from semantic refinement. On a 1,094-sample evaluation set, OmniDirector outperforms CamCloneMaster, Seedance2.0, and LTX-LoRA, reaching 2.64° relative rotation error, 16.84° relative translation error, 83.18% rotation precision, 72.74% translation precision, 96.52% transition temporal precision, 83.79% transition semantic precision, and only 0.51% frame leakage and 3.38% shot leakage. The paper’s key novelty is that this camera-grid representation also enables zero-shot camera understanding from raw RGB or Canny-edge inputs, showing that the model can generalize beyond explicitly paired camera data.
Original abstract
Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations that fail to handle multi-shot generation or synthesize cross-paired data, which suffer from data scarcity, resulting in poor performance in complicated camera motion cloning. To address these issues, we introduce a general camera motion representation that encodes cameras as grid motion videos. This camera grid represents the camera parameters visually and supports the integration of diverse trajectories for multi-shot video generation. Building upon this, we propose OmniDirector, a unified framework trained on a million-scale camera grid-video pairs that coordinates characters, actions, and cameras to provide director-level control for multimodal diffusion transformers. Furthermore, we design a novel hierarchical prompt expansion agent that harmoniously integrates different control signals by systematically describing camera motion and visual content through understanding signal relationships. Extensive experiments demonstrate the superior performance and outstanding controllability of our framework. Project page: https://ymlinfeng.github.io/OmniDirector.github.io/
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.