MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
AuthorsGal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
Resources
MV-Forcing lets a diffusion model generate long, consistent videos from multiple viewpoints by using 3D geometry to guide each new view and time step.
Key results
Synthetic training and evaluation dataset size
Synchronized cameras per synthetic scene
Evaluation configuration for long multi-view generation
Evaluation configuration for long multi-view generation
Camera accuracy on synthetic long-sequence evaluation
Camera accuracy on synthetic long-sequence evaluation
What the paper found
MV-Forcing, from researchers at the Hebrew University of Jerusalem and Cornell University, tackles a gap that existing video diffusion systems such as SynCamMaster and Self-Forcing leave open: generating long videos that stay consistent across multiple camera viewpoints. The key idea is to combine temporal autoregression with view-sequential autoregression, while using CUT3R as a 4D geometric bridge that reconstructs a persistent scene state from previously generated frames and renders a viewpoint-specific geometric prior for the next view. The student model is distilled from a bidirectional SynCamMaster teacher with Distribution Matching Distillation, then trained with spatio-temporal self-forcing and a joint denoising regime so the first view can be generated from text and later views can be generated autoregressively without a fixed temporal window. On the SynCamVideo benchmark, which contains 3,400 synthetic scenes captured from 10 synchronized cameras, MV-Forcing reaches strong long-horizon performance at 3 views and 162 frames, improving camera accuracy to 3.72 rotation error and 8.41 translation error, while raising cross-view synchronization to 243.63 matching pixels and 89.88 CLIP-V in the synthetic setting. Scaling is notably stable: performance remains close to constant from 2 to 5 views, and across 81 to 648 frames the method preserves camera metrics with only mild temporal degradation. Runtime is also practical, with a 4-step causal student running in about 70 seconds for 2 views at 81 frames and using 23 GB peak VRAM, roughly a 5× speedup over the 50-step SynCamMaster baseline.
Original abstract
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.