NTH

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

AuthorsYuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu, Junsong Chen, Tian Ye, Haozhe Liu, Enze Xie, Song Han

June 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes real-time video editing practical on a consumer GPU by pairing a diffusion-transformer model with a new consistency training trick and aggressive Blackwell-specific optimization.

Key results

2B
Model size

SANA-Streaming is described as a 2B hybrid diffusion transformer.

5.56 GB
VRAM

The hybrid architecture requires 5.56 GB VRAM for long video generation.

24
End-to-end FPS

The system achieves 24 end-to-end FPS for 1280×704 real-time streaming video editing on a single RTX 5090.

58
DiT FPS

The DiT core runs at 58 FPS on an RTX 5090.

2.62
OpenVE-Bench average score

On OpenVE-Bench, SANA-Streaming reports a 2.62 average score across the five spatial edit categories.

What the paper found

SANA-Streaming, from NVIDIA and MIT researchers led by Song Han, tackles a hard limitation in video-to-video editing: maintaining source fidelity and temporal coherence while streaming minute-long, high-resolution outputs in real time on a consumer GPU. The core model is a 2B hybrid diffusion transformer that interleaves 15 Gated DeltaNet linear-attention blocks with 5 softmax-attention blocks, combining constant-memory global recurrence with local window-plus-sink refinement; this hybrid design cuts VRAM to 5.56 GB and makes the DiT core run at 58 FPS. To train without paired long edited videos, the paper introduces Cycle-Reverse Regularization, where a forward edit is followed by a reverse prompt that reconstructs the original source chunk using a flow-matching loss, improving long-range consistency and reducing drift. The system also distills a causal VAE decoder from the LTX2 VAE so streaming decoding does not depend on future frames, and it pairs this with Triton fused GDN kernels plus mixed-precision quantization on NVIDIA Blackwell, yielding a 1.59× DiT speedup over BF16 and 24 end-to-end FPS at 1280×704 on a single RTX 5090. On OpenVE-Bench, the method reports a 2.62 average score across the five spatial edit categories, outperforming prior systems such as OpenVE-Edit and Lucy-Edit while being far faster. The paper’s novelty is the joint treatment of architecture, training, and GPU systems for real-time interactive video editing.

Original abstract

Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements for temporal consistency and inference throughput. In this paper, we present SANA-Streaming, a system-algorithm co-designed framework for high-resolution, real-time streaming video editing on consumer GPUs, with the following three core designs: (1) Hybrid Diffusion Transformer architecture introduces softmax attention in part of the blocks to improve local modeling capabilities while preserving the efficiency of linear layers. (2) Cycle-Reverse Regularization is a novel training strategy that enforces semantic consistency by predicting source frames from generated content via flow matching, improving temporal consistency without requiring paired long edited videos. (3) Efficient System Co-design combines fused GDN kernels and Mixed-Precision Quantization (MPQ) optimized for the NVIDIA Blackwell (RTX 5090) architecture. By profiling real-world throughput, our MPQ maximizes Tensor Core utilization while maintaining generation quality. The resulting system achieves real-time 1280 x 704 resolution editing at 24 end-to-end FPS on a single RTX 5090 GPU, with the DiT core running at 58 FPS. Experimental results demonstrate that our co-design approach significantly outperforms existing SOTA methods in both temporal coherence and system throughput.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis