NTH

4DStreamCtrl: Interactive Video Generation with Online 4D Control

AuthorsShiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu

September 2, 2026 3 min read
Watch on YouTube
The one-line take

4DStreamCtrl aims to make video diffusion interactive by generating long, realistic streams while jointly controlling camera motion, object movement, and depth in real time.

Key results

0.4M
OpenVidHD-Motion3D clips

Retained training clips with 3D point tracks and camera parameters

5B
Wan2.2 TI2V-5B backbone

Pretrained diffusion transformer used by the teacher and student

4
Student denoising steps

Causal streaming student distilled from the bidirectional teacher

20.6
Streaming throughput

FPS achieved for 480p video on one high-end GPU

5.29
DAVIS teacher EPE

Motion-control error on the DAVIS validation set

350
Long-video coherence

Frames sustained with temporally coherent streaming generation

What the paper found

4DStreamCtrl targets a limitation in video generators such as Wan and Tencent HunyuanVideo: text and 2D trajectories do not precisely distinguish object motion from camera motion, while existing 3D methods are offline. It unifies camera motion, object trajectories, and depth as 3D point tracks, encoded by a lightweight Geometric Motion Head and injected into the Wan2.2 TI2V-5B diffusion transformer. Training uses OpenVidHD-Motion3D, mined from OpenVid-1M with SpatialTrackerV2 and retaining roughly 0.4M clips. The same interface supports joint camera-object control, depth-aware editing, and motion transfer; for appearance changes, the system can restyle a first frame with Stable Diffusion XL while preserving source 3D dynamics. To enable interaction, self-forcing and distribution matching distillation convert a bidirectional 50-step teacher into a causal student requiring 4 denoising steps, using block-wise causal attention, an attention sink, and rolling KV-cache reuse. The student reaches 20.6 FPS for 480p video on one high-end GPU, with memory independent of video length, and remains coherent for 350 frames. On the DAVIS validation set, the teacher achieves an EPE of 5.29, compared with 7.86 for a matched Wan2.2 2D-track baseline, while the causal student records EPE 5.48 at interactive speed. The result is a unified system for online 4D-controlled streaming generation rather than offline, fixed-length synthesis.

Original abstract

Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing approach captures only part of this: camera-parameter methods steer the viewpoint but cannot move objects, 2D-trajectory methods act in the image plane and ignore depth and occlusion, and recent 3D methods add geometry but run only offline at a fixed length. In particular, none combines 3D-consistent control of both camera and objects with real-time, streaming generation. Here we show that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass. To learn this interface at scale, we mine in-the-wild video for 3D motion supervision, yielding OpenVidHD-Motion3D, and encode it with a lightweight Geometric Motion Head that plugs into a pretrained video diffusion model. Because this encoder is temporally separable, we distill the model into a causal streaming student that generates arbitrarily long video in four denoising steps at memory independent of length. This unified design surpasses prior camera-only, 2D, and offline-3D methods in motion-control precision while covering modalities they address only in isolation. 4DStreamCtrl runs at 20 FPS on a single high-end GPU for 480p video and stays temporally coherent over hundreds of frames, enabling, to our knowledge, interactive 4D-controllable streaming generation for the first time. More broadly, grounding generation in explicit 3D geometry with efficient causal inference points toward interactive world models with closed-loop spatiotemporal control, from controllable simulators to real-time visual imagination for embodied agents.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis