NTH

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

AuthorsKaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, Huayu Chen, Jianfei Chen, Chen-Hsuan Lin, Ming-Yu Liu, Jun Zhu, Qianli Ma

June 25, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to make streaming video and interactive world-model diffusion much faster and more practical by unifying two training styles into one distillation recipe.

Key results

84.63
VBench-T2V score

Causal-rCM distilled Wan2.1-1.3B achieves this score with 1 or 2 sampling steps.

10x
speedup

Teacher-forcing sCM/MeanFlow converges 10× faster than discrete-time CMs.

83.35
Wan2.1-14B VBench-T2V

Bidirectional teacher baseline on the same benchmark.

What the paper found

Causal-rCM, developed by Tsinghua University, UT Austin, and NVIDIA, extends rCM from standard diffusion distillation to autoregressive video diffusion by pairing teacher-forcing consistency models with self-forcing distribution matching in a single causal recipe. The paper argues that teacher-forcing provides the stable forward-divergence initialization needed to combat exposure bias, while self-forcing refines the student on its own rollouts to match the inference-time distribution. A key technical contribution is the first teacher-forcing implementation of continuous-time consistency models, including sCM and MeanFlow, for autoregressive video, enabled by a custom-mask FlashAttention-2 JVP kernel; this yields 10× faster convergence than discrete-time consistency models. On Wan2.1-1.3B text-to-video, the distilled model reaches a VBench-T2V score of 84.63 with only 1 or 2 sampling steps, outperforming the 14B bidirectional baseline at 83.35 and prior streaming methods such as Self-Forcing, LongLive, Causal Forcing, and AnyFlow. The recipe is staged in three phases: TF adaptation, TF-CM distillation, and SF-DMD refinement, and it also extends to NVIDIA Cosmos 3, where the GEN vision stream is converted into temporal-causal autoregressive diffusion for interactive world modeling with action-conditioned control.

Original abstract

Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models. In this work, we extend rCM, an advanced diffusion distillation framework, to autoregressive video diffusion. The core philosophy of rCM lies in the complementarity between forward and reverse divergences, represented by consistency models (CMs) and distribution matching distillation (DMD), respectively, in diffusion distillation. This philosophy naturally carries over to the autoregressive setting, where teacher-forcing (TF) provides an offline, forward-divergence causal training paradigm, while self-forcing (SF) corresponds to an on-policy, reverse-divergence refinement. Our contributions are: (1) through extensive experiments, we show that teacher-forcing CM is currently the best complement to self-forcing DMD as an initialization strategy (2) we present the first implementation of teacher-forcing-based continuous-time CMs (e.g., sCM/MeanFlow) for autoregressive video diffusion, enabled by our custom-mask FlashAttention-2 JVP kernel, achieving 10$\times$ faster convergence compared to discrete-time CMs (dCMs) (3) we introduce Causal-rCM, a leading, unified, and scalable algorithm-infrastructure open recipe for diffusion distillation and causal training (4) we achieve state-of-the-art streaming video generation performance in both frame-wise and chunk-wise settings, using only synthetic data for training. Notably, our distilled 2-step causal Wan2.1-1.3B model achieves a VBench-T2V score of 84.63 with only 1 or 2 sampling steps. We further apply Causal-rCM to Cosmos 3, an advanced omnimodal world foundation model for physical AI with action-conditioned generation capability, enabling an interactive world model.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis