NTH

4Director: Controlling Video World Models with Rigid 3D Geometry

AuthorsWei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

AffiliationsOct 2026 Camera Control Input Image Figure 1: Camera and object control from a single image. Top: the image is lifted into a back- ground point cloud and complete rigid 3D geometry for each object, in

October 5, 2026 2 min read
Watch on YouTube
The one-line take

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Key results

20774
RealCOD-Rigid clips

Annotated training clips containing reconstructed rigid 3D scenes.

3.049B
Motion Adapter parameters

Trainable adapter capacity built on Wan2.1-VACE-14B.

44.1
FID

Best visual-quality score in the joint camera and object-control comparison.

370.4
FVD

Best video-distribution quality score in the comparison.

3.65
Camera rotation error

Mean rotation trajectory error in degrees.

60.4
Identity-Gated IoU

Object-placement score gated by preservation of object identity.

What the paper found

4Director is a video world model that replaces ambiguous 2D motion cues with an explicit 4D scene: each marked object is reconstructed from one image as a complete canonical mesh, moved by one rigid SE(3) transformation per frame, and rendered alongside a background point cloud and camera trajectory. The system uses MoGe-2 for depth, SAM 2 for segmentation, and Pixal3D for complete object geometry, then feeds the resulting depth video into a 3.049B-parameter Motion Adapter built on Wan2.1-VACE-14B. The adapter preserves prescribed camera motion, object translation, rotation, occlusion, and identity while synthesizing appearance, lighting, background completion, and non-rigid dynamics; Qwen3-VL supplies identity judgments for the proposed Identity-Gated IoU metric. Training uses the RealCOD-Rigid dataset of 20774 annotated clips. On 100 held-out clips, 4Director achieves FID 44.1, FVD 370.4, camera rotation error 3.65 degrees, and IG-IoU 60.4, outperforming MotionCtrl, Perception-as-Control, VerseCrafter, and SymphoMotion. The main limitation is that one rigid transform cannot prescribe articulated motion, so actions such as dancing or limb movement remain under generator control.

Original abstract

Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
02World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis
03World Model

HappyWorld-Bench

Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu

HappyWorld-Bench tests whether AI-generated worlds remain coherent, editable, and responsive when agents explore and act within them.

Read analysis