Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
AuthorsXinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang
Resources
Flex-Forcing lets one video diffusion model switch between fast autoregressive generation and high-quality bidirectional generation, aiming to get the best of both worlds.
Key results
Base video model and teacher-scale setting used for Flex-Forcing experiments
Best-performing 5-second video inference speed on GB200
Best-performing VBench Total Quality score for 5-second videos
Overall long-video benchmark score on 30-second videos
What the paper found
Flex-Forcing, from NVIDIA Research with Xinchao Wang and Arash Vahdat among the authors, reframes video diffusion as a single model that can switch at test time between bidirectional and autoregressive generation. Built on Wan2.1-T2V-1.3B and trained with a 1.3B teacher, it introduces flexible chunking over both frames and denoising timesteps, so early high-noise steps can use coarse global chunks while later steps use finer local chunks. A key technical addition is K-Projection, a timestep-conditioned key-cache alignment that resolves the signal-to-noise mismatch when causal past tokens and non-causal future tokens coexist in attention. On 5-second video generation, the paper reports that Flex-Forcing reaches 85.07 VBench Total Quality and 86.33 Semantic at 25.8 FPS with 5 denoising steps, while a faster setting reaches 29.4 FPS at 84.63 Total Quality and 85.89 Semantic. In the 2-step regime, it still attains 84.20 Total Quality and 85.27 Semantic at 24.9 FPS. For 30-second video on VBench-Long, it improves total score to 84.01 and raises FPS to 24.96, outperforming Self-Forcing and Infinity-RoPE on long-horizon coherence, motion smoothness, and dynamic degree. The model also enables any-order, any-step autoregressive editing, allowing localized video edits without regenerating the full sequence, which is a practical step toward controllable long-video synthesis.
Original abstract
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive models offer efficient and streaming generation at the cost of long-range consistency and exposure bias. We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. The core idea is a flexible chunking mechanism jointly defined over the temporal axis and denoising steps. This design allows the model to (1) perform flexible chunking according to different device budgets, (2) perform bidirectional inference across chunks for global structure planning, while generating frames autoregressively within each chunk for efficient and fine-grained synthesis, and (3) perform any-order, any-timestep autoregressive generation without the strict causal constraint. Extensive experiments on multiple video generation benchmarks demonstrate that Flex-Forcing achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule, while offering faster inference.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.