Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
AuthorsKaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, Huayu Chen, Jianfei Chen, Chen-Hsuan Lin, Ming-Yu Liu, Jun Zhu, Qianli Ma
Resources
This paper shows how to make streaming video and interactive world-model diffusion much faster and more practical by unifying two training styles into one distillation recipe.
Key results
Causal-rCM distilled Wan2.1-1.3B achieves this score with 1 or 2 sampling steps.
Teacher-forcing sCM/MeanFlow converges 10× faster than discrete-time CMs.
Bidirectional teacher baseline on the same benchmark.
What the paper found
Causal-rCM, developed by Tsinghua University, UT Austin, and NVIDIA, extends rCM from standard diffusion distillation to autoregressive video diffusion by pairing teacher-forcing consistency models with self-forcing distribution matching in a single causal recipe. The paper argues that teacher-forcing provides the stable forward-divergence initialization needed to combat exposure bias, while self-forcing refines the student on its own rollouts to match the inference-time distribution. A key technical contribution is the first teacher-forcing implementation of continuous-time consistency models, including sCM and MeanFlow, for autoregressive video, enabled by a custom-mask FlashAttention-2 JVP kernel; this yields 10× faster convergence than discrete-time consistency models. On Wan2.1-1.3B text-to-video, the distilled model reaches a VBench-T2V score of 84.63 with only 1 or 2 sampling steps, outperforming the 14B bidirectional baseline at 83.35 and prior streaming methods such as Self-Forcing, LongLive, Causal Forcing, and AnyFlow. The recipe is staged in three phases: TF adaptation, TF-CM distillation, and SF-DMD refinement, and it also extends to NVIDIA Cosmos 3, where the GEN vision stream is converted into temporal-causal autoregressive diffusion for interactive world modeling with action-conditioned control.
Original abstract
Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models. In this work, we extend rCM, an advanced diffusion distillation framework, to autoregressive video diffusion. The core philosophy of rCM lies in the complementarity between forward and reverse divergences, represented by consistency models (CMs) and distribution matching distillation (DMD), respectively, in diffusion distillation. This philosophy naturally carries over to the autoregressive setting, where teacher-forcing (TF) provides an offline, forward-divergence causal training paradigm, while self-forcing (SF) corresponds to an on-policy, reverse-divergence refinement. Our contributions are: (1) through extensive experiments, we show that teacher-forcing CM is currently the best complement to self-forcing DMD as an initialization strategy (2) we present the first implementation of teacher-forcing-based continuous-time CMs (e.g., sCM/MeanFlow) for autoregressive video diffusion, enabled by our custom-mask FlashAttention-2 JVP kernel, achieving 10$\times$ faster convergence compared to discrete-time CMs (dCMs) (3) we introduce Causal-rCM, a leading, unified, and scalable algorithm-infrastructure open recipe for diffusion distillation and causal training (4) we achieve state-of-the-art streaming video generation performance in both frame-wise and chunk-wise settings, using only synthetic data for training. Notably, our distilled 2-step causal Wan2.1-1.3B model achieves a VBench-T2V score of 84.63 with only 1 or 2 sampling steps. We further apply Causal-rCM to Cosmos 3, an advanced omnimodal world foundation model for physical AI with action-conditioned generation capability, enabling an interactive world model.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.