NTH

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

AuthorsGal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim

July 12, 2026 2 min read
Watch on YouTube
The one-line take

MV-Forcing lets a diffusion model generate long, consistent videos from multiple viewpoints by using 3D geometry to guide each new view and time step.

Key results

3400
SynCamVideo scenes

Synthetic training and evaluation dataset size

10
SynCamVideo cameras

Synchronized cameras per synthetic scene

3
Long-sequence views

Evaluation configuration for long multi-view generation

162
Long-sequence frames

Evaluation configuration for long multi-view generation

3.72
Rotation error

Camera accuracy on synthetic long-sequence evaluation

8.41
Translation error

Camera accuracy on synthetic long-sequence evaluation

What the paper found

MV-Forcing, from researchers at the Hebrew University of Jerusalem and Cornell University, tackles a gap that existing video diffusion systems such as SynCamMaster and Self-Forcing leave open: generating long videos that stay consistent across multiple camera viewpoints. The key idea is to combine temporal autoregression with view-sequential autoregression, while using CUT3R as a 4D geometric bridge that reconstructs a persistent scene state from previously generated frames and renders a viewpoint-specific geometric prior for the next view. The student model is distilled from a bidirectional SynCamMaster teacher with Distribution Matching Distillation, then trained with spatio-temporal self-forcing and a joint denoising regime so the first view can be generated from text and later views can be generated autoregressively without a fixed temporal window. On the SynCamVideo benchmark, which contains 3,400 synthetic scenes captured from 10 synchronized cameras, MV-Forcing reaches strong long-horizon performance at 3 views and 162 frames, improving camera accuracy to 3.72 rotation error and 8.41 translation error, while raising cross-view synchronization to 243.63 matching pixels and 89.88 CLIP-V in the synthetic setting. Scaling is notably stable: performance remains close to constant from 2 to 5 views, and across 81 to 648 frames the method preserves camera metrics with only mild temporal degradation. Runtime is also practical, with a 4-step causal student running in about 70 seconds for 2 views at 81 frames and using 23 GB peak VRAM, roughly a 5× speedup over the 50-step SynCamMaster baseline.

Original abstract

Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis