NTH

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

AuthorsZekun Li, Xiaoyan Cong, Hongyu Li, Zhiyang Dou, Chuan Guo, Abhay Mittal, Sizhe An, Srinath Sridhar

August 2, 2026 2 min read
Watch on YouTube
The one-line take

Ms.Forcing accelerates real-time video diffusion by using coarser representations for noisier frames while preserving or improving generation quality.

Key results

45%
Active-window token reduction

Reduction from Multi-Scale Patchification.

12.9%
Post-MSP attention reduction

Additional dot-product attention reduction from Multi-Scale Self-Attention.

22.84
Ms. Forcing throughput

Frames per second on one H200 GPU at 832×480 resolution.

39.6%
Throughput gain over Rolling Forcing

Relative inference-speed improvement.

0.29
Five-second VBench quality improvement

Quality-score improvement over Rolling Forcing.

1.700
60-second quality drift

Final drift score, reduced from 2.227.

What the paper found

Ms. Forcing, developed by researchers at Brown University, MIT, and Meta, accelerates streaming video generation by redesigning Rolling Forcing’s mixed-noise denoising window. Built on the Wan2.1-T2V-1.3B backbone, the method uses Multi-Scale Patchification, assigning coarse spatial patches to high-noise states and fine patches to near-clean states, reducing active-window tokens by 45%. Multi-Scale Self-Attention further subsamples visible non-sink keys and values according to each query’s spatial scale, cutting post-patchification dot-product attention by 12.9% while preserving a full-resolution attention sink and a static, hardware-friendly computation graph. The third component, Homogeneous-Noise-Level Distribution Matching Distillation, assembles training videos from predictions sharing the same source noise level, reducing the mismatch between training rollouts and inference. At 832×480 resolution, Ms. Forcing reaches 22.84 FPS on one H200 GPU, 39.6% faster than Rolling Forcing. On five-second VBench videos, it improves quality and semantic scores over Rolling Forcing by 0.29 and 1.25, respectively; on 60-second generation, it reduces quality drift from 2.227 to 1.700. The experiments show that noise-matched spatial computation can improve both throughput and long-horizon stability, although KV-cache updates remain a latency bottleneck.

Original abstract

Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional next-frame generation hinder real-time deployment. Recent rolling-window methods pipeline denoising across multiple consecutive frames at different noise levels, improving throughput and long-horizon stability. However, they tokenize every state at the same fine spatial granularity, leaving substantial noise-dependent redundancy in the joint denoising window. We propose Ms.Forcing, an efficient streaming video generation paradigm that adapts spatial granularity to each state's noise level. Its Multi-Scale Patchification (MSP) assigns coarser patches to noisier states, reducing the active-window token count by 45%, while Multi-Scale Self-Attention (MSSA) matches the density of visible non-sink keys and values to each query scale to further reduce attention cost. Because both schedules are fixed by window position, Ms.Forcing retains a static, hardware-friendly computation graph. We further introduce Homogeneous-Noise-Level DMD (H-DMD), which assembles each fake video from clean predictions sharing the same source noise level, thereby reducing the mismatch between DMD training sequences and inference-time rollouts. The multi-scale design helps offset the additional training cost of backpropagating through overlapping windows. We include both quantitative and qualitative experiments to show that Ms.Forcing reaches 22.84 FPS on a single H200 GPU, 39.6% faster than Rolling Forcing, while significantly improving VBench scores in both short video and long video generation setting.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis