NTH

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

AuthorsXinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang

July 11, 2026 2 min read
Watch on YouTube
The one-line take

Flex-Forcing lets one video diffusion model switch between fast autoregressive generation and high-quality bidirectional generation, aiming to get the best of both worlds.

Key results

1.3B
Model size

Base video model and teacher-scale setting used for Flex-Forcing experiments

25.8
5s FPS

Best-performing 5-second video inference speed on GB200

85.07
5s Total Quality

Best-performing VBench Total Quality score for 5-second videos

84.01
30s VBench-Long Total Score

Overall long-video benchmark score on 30-second videos

What the paper found

Flex-Forcing, from NVIDIA Research with Xinchao Wang and Arash Vahdat among the authors, reframes video diffusion as a single model that can switch at test time between bidirectional and autoregressive generation. Built on Wan2.1-T2V-1.3B and trained with a 1.3B teacher, it introduces flexible chunking over both frames and denoising timesteps, so early high-noise steps can use coarse global chunks while later steps use finer local chunks. A key technical addition is K-Projection, a timestep-conditioned key-cache alignment that resolves the signal-to-noise mismatch when causal past tokens and non-causal future tokens coexist in attention. On 5-second video generation, the paper reports that Flex-Forcing reaches 85.07 VBench Total Quality and 86.33 Semantic at 25.8 FPS with 5 denoising steps, while a faster setting reaches 29.4 FPS at 84.63 Total Quality and 85.89 Semantic. In the 2-step regime, it still attains 84.20 Total Quality and 85.27 Semantic at 24.9 FPS. For 30-second video on VBench-Long, it improves total score to 84.01 and raises FPS to 24.96, outperforming Self-Forcing and Infinity-RoPE on long-horizon coherence, motion smoothness, and dynamic degree. The model also enables any-order, any-step autoregressive editing, allowing localized video edits without regenerating the full sequence, which is a practical step toward controllable long-video synthesis.

Original abstract

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive models offer efficient and streaming generation at the cost of long-range consistency and exposure bias. We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. The core idea is a flexible chunking mechanism jointly defined over the temporal axis and denoising steps. This design allows the model to (1) perform flexible chunking according to different device budgets, (2) perform bidirectional inference across chunks for global structure planning, while generating frames autoregressively within each chunk for efficient and fine-grained synthesis, and (3) perform any-order, any-timestep autoregressive generation without the strict causal constraint. Extensive experiments on multiple video generation benchmarks demonstrate that Flex-Forcing achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule, while offering faster inference.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis