4DStreamCtrl: Interactive Video Generation with Online 4D Control
AuthorsShiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
Resources
4DStreamCtrl aims to make video diffusion interactive by generating long, realistic streams while jointly controlling camera motion, object movement, and depth in real time.
Key results
Retained training clips with 3D point tracks and camera parameters
Pretrained diffusion transformer used by the teacher and student
Causal streaming student distilled from the bidirectional teacher
FPS achieved for 480p video on one high-end GPU
Motion-control error on the DAVIS validation set
Frames sustained with temporally coherent streaming generation
What the paper found
4DStreamCtrl targets a limitation in video generators such as Wan and Tencent HunyuanVideo: text and 2D trajectories do not precisely distinguish object motion from camera motion, while existing 3D methods are offline. It unifies camera motion, object trajectories, and depth as 3D point tracks, encoded by a lightweight Geometric Motion Head and injected into the Wan2.2 TI2V-5B diffusion transformer. Training uses OpenVidHD-Motion3D, mined from OpenVid-1M with SpatialTrackerV2 and retaining roughly 0.4M clips. The same interface supports joint camera-object control, depth-aware editing, and motion transfer; for appearance changes, the system can restyle a first frame with Stable Diffusion XL while preserving source 3D dynamics. To enable interaction, self-forcing and distribution matching distillation convert a bidirectional 50-step teacher into a causal student requiring 4 denoising steps, using block-wise causal attention, an attention sink, and rolling KV-cache reuse. The student reaches 20.6 FPS for 480p video on one high-end GPU, with memory independent of video length, and remains coherent for 350 frames. On the DAVIS validation set, the teacher achieves an EPE of 5.29, compared with 7.86 for a matched Wan2.2 2D-track baseline, while the causal student records EPE 5.48 at interactive speed. The result is a unified system for online 4D-controlled streaming generation rather than offline, fixed-length synthesis.
Original abstract
Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing approach captures only part of this: camera-parameter methods steer the viewpoint but cannot move objects, 2D-trajectory methods act in the image plane and ignore depth and occlusion, and recent 3D methods add geometry but run only offline at a fixed length. In particular, none combines 3D-consistent control of both camera and objects with real-time, streaming generation. Here we show that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass. To learn this interface at scale, we mine in-the-wild video for 3D motion supervision, yielding OpenVidHD-Motion3D, and encode it with a lightweight Geometric Motion Head that plugs into a pretrained video diffusion model. Because this encoder is temporally separable, we distill the model into a causal streaming student that generates arbitrarily long video in four denoising steps at memory independent of length. This unified design surpasses prior camera-only, 2D, and offline-3D methods in motion-control precision while covering modalities they address only in isolation. 4DStreamCtrl runs at 20 FPS on a single high-end GPU for 480p video and stays temporally coherent over hundreds of frames, enabling, to our knowledge, interactive 4D-controllable streaming generation for the first time. More broadly, grounding generation in explicit 3D geometry with efficient causal inference points toward interactive world models with closed-loop spatiotemporal control, from controllable simulators to real-time visual imagination for embodied agents.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.