SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
AuthorsJunsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie
Resources
SANA-Video 2.0 combines mostly linear attention with occasional softmax anchors to make high-quality long-video generation dramatically faster on a single GPU.
Key results
The selected hybrid layout uses 25% gated-softmax anchors and 75% linear-attention layers.
Block Attention Residuals increase effective rank in deep linear-attention layers.
The 5B model’s VBench score after 40-step sampling.
Time to generate an 81-frame 480p video on one NVIDIA H100.
Speedup over matched full-softmax attention at a 720p/60s tensor shape.
End-to-end acceleration from kernel optimization, diffusion caching, and sparse attention.
What the paper found
NVIDIA researchers introduce SANA-Video 2.0, a video diffusion transformer available at 5B and 14B scales that targets the quadratic cost of full-softmax attention. Its hybrid architecture uses 75% gated bidirectional linear-attention layers and 25% periodic gated-softmax anchors in a 3:1 layout, preserving long-sequence scaling while periodically restoring full-rank token interactions. Block Attention Residuals, or AttnRes, route completed eight-layer block summaries into later layers so linear attention can reuse information refreshed by the anchors; this increases deep-layer effective rank by 12%. The model is trained from scratch with flow matching, a resolution-and-duration curriculum, structured video curation, Self-Flow, Diffusion-DPO, and ReFL, rather than converting a pretrained softmax model. On the VBench benchmark, the 5B model scores 84.30 after 40-step sampling, generating an 81-frame 480p video in 13.2s on a single NVIDIA H100. At a 720p/60s tensor shape, its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline, with the advantage increasing as sequences lengthen. NVIDIA’s Sol-Engine deployment stack adds kernel fusion, diffusion caching, and sparse anchor attention, producing a separately measured 3.58x end-to-end speedup and enabling efficient high-resolution generation. The design draws on hybrid-attention patterns explored in Qwen3-Next and Kimi-Linear, but reselects the attention ratio for video and adapts cross-depth residual routing to bidirectional diffusion.
Original abstract
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.