NTH

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

AuthorsJiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen Liu, Pengfei Wan

June 16, 2026 2 min read
Watch on YouTube
The one-line take

OmniDirector makes video generation follow complex multi-shot camera moves by turning camera paths into visual motion grids and training a large controllable diffusion system on them.

Key results

1.8M
training video corpus

internet videos used to build the camera grid–video dataset

30%
self-reconstruction share

portion of training samples using camera-grid reconstruction

1094
evaluation set size

curated validation samples

2.64
relative rotation error

OmniDirector camera accuracy

72.74%
translation precision

OmniDirector T-Pre on the evaluation set

0.51%
frame leakage

reference-video leakage at frame level

What the paper found

OmniDirector, developed by the Kuaishou Technology Kling team with Tsinghua University and Peking University collaborators, tackles general multi-shot camera cloning by replacing brittle camera-parameter injection and scarce cross-paired data with a visual camera grid: reference videos are decomposed into camera poses, rendered as motion in an empty 3D scene, and fed into a multimodal diffusion transformer as a spatiotemporal conditioning stream. The method is trained on a million-scale camera grid–video dataset built from 1.8M internet videos, then fine-tuned with a self-reconstruction objective on 30% of samples to force the model to learn geometry rather than appearance leakage. At inference, a hierarchical Prompt Expansion agent uses Qwen3-VL to generate inter-shot and intra-shot camera descriptions and fuse them with the user prompt and reference image, while adaptive classifier-free guidance separates camera structure from semantic refinement. On a 1,094-sample evaluation set, OmniDirector outperforms CamCloneMaster, Seedance2.0, and LTX-LoRA, reaching 2.64° relative rotation error, 16.84° relative translation error, 83.18% rotation precision, 72.74% translation precision, 96.52% transition temporal precision, 83.79% transition semantic precision, and only 0.51% frame leakage and 3.38% shot leakage. The paper’s key novelty is that this camera-grid representation also enables zero-shot camera understanding from raw RGB or Canny-edge inputs, showing that the model can generalize beyond explicitly paired camera data.

Original abstract

Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations that fail to handle multi-shot generation or synthesize cross-paired data, which suffer from data scarcity, resulting in poor performance in complicated camera motion cloning. To address these issues, we introduce a general camera motion representation that encodes cameras as grid motion videos. This camera grid represents the camera parameters visually and supports the integration of diverse trajectories for multi-shot video generation. Building upon this, we propose OmniDirector, a unified framework trained on a million-scale camera grid-video pairs that coordinates characters, actions, and cameras to provide director-level control for multimodal diffusion transformers. Furthermore, we design a novel hierarchical prompt expansion agent that harmoniously integrates different control signals by systematically describing camera motion and visual content through understanding signal relationships. Extensive experiments demonstrate the superior performance and outstanding controllability of our framework. Project page: https://ymlinfeng.github.io/OmniDirector.github.io/

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis