SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
AuthorsYuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu, Junsong Chen, Tian Ye, Haozhe Liu, Enze Xie, Song Han
Resources
This paper makes real-time video editing practical on a consumer GPU by pairing a diffusion-transformer model with a new consistency training trick and aggressive Blackwell-specific optimization.
Key results
SANA-Streaming is described as a 2B hybrid diffusion transformer.
The hybrid architecture requires 5.56 GB VRAM for long video generation.
The system achieves 24 end-to-end FPS for 1280×704 real-time streaming video editing on a single RTX 5090.
The DiT core runs at 58 FPS on an RTX 5090.
On OpenVE-Bench, SANA-Streaming reports a 2.62 average score across the five spatial edit categories.
What the paper found
SANA-Streaming, from NVIDIA and MIT researchers led by Song Han, tackles a hard limitation in video-to-video editing: maintaining source fidelity and temporal coherence while streaming minute-long, high-resolution outputs in real time on a consumer GPU. The core model is a 2B hybrid diffusion transformer that interleaves 15 Gated DeltaNet linear-attention blocks with 5 softmax-attention blocks, combining constant-memory global recurrence with local window-plus-sink refinement; this hybrid design cuts VRAM to 5.56 GB and makes the DiT core run at 58 FPS. To train without paired long edited videos, the paper introduces Cycle-Reverse Regularization, where a forward edit is followed by a reverse prompt that reconstructs the original source chunk using a flow-matching loss, improving long-range consistency and reducing drift. The system also distills a causal VAE decoder from the LTX2 VAE so streaming decoding does not depend on future frames, and it pairs this with Triton fused GDN kernels plus mixed-precision quantization on NVIDIA Blackwell, yielding a 1.59× DiT speedup over BF16 and 24 end-to-end FPS at 1280×704 on a single RTX 5090. On OpenVE-Bench, the method reports a 2.62 average score across the five spatial edit categories, outperforming prior systems such as OpenVE-Edit and Lucy-Edit while being far faster. The paper’s novelty is the joint treatment of architecture, training, and GPU systems for real-time interactive video editing.
Original abstract
Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements for temporal consistency and inference throughput. In this paper, we present SANA-Streaming, a system-algorithm co-designed framework for high-resolution, real-time streaming video editing on consumer GPUs, with the following three core designs: (1) Hybrid Diffusion Transformer architecture introduces softmax attention in part of the blocks to improve local modeling capabilities while preserving the efficiency of linear layers. (2) Cycle-Reverse Regularization is a novel training strategy that enforces semantic consistency by predicting source frames from generated content via flow matching, improving temporal consistency without requiring paired long edited videos. (3) Efficient System Co-design combines fused GDN kernels and Mixed-Precision Quantization (MPQ) optimized for the NVIDIA Blackwell (RTX 5090) architecture. By profiling real-world throughput, our MPQ maximizes Tensor Core utilization while maintaining generation quality. The resulting system achieves real-time 1280 x 704 resolution editing at 24 end-to-end FPS on a single RTX 5090 GPU, with the DiT core running at 58 FPS. Experimental results demonstrate that our co-design approach significantly outperforms existing SOTA methods in both temporal coherence and system throughput.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.