VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
AuthorsHidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay, Hoda Eldardiry, Pinar Yanardag
Resources
VideoMLA makes minute-scale autoregressive video diffusion much cheaper in memory by compressing the KV cache with a low-rank latent representation, enabling longer video rollouts with better throughput and strong quality.
Key results
VideoMLA reduces per-token cached KV state from 3072 dense scalars to 224 scalars per layer.
Default VideoMLA setting stores 224 scalars per cached token per layer using a 192-dim content latent plus 32 RoPE channels.
Wan2.1-T2V-1.3B dense per-token, per-layer KV cache stores 2 × 12 × 128 = 3072 scalars.
VideoMLA achieves the best overall VBench score at 60 seconds among evaluated methods.
VideoMLA reaches 23.96 FPS on B200 in the chunk-wise autoregressive comparison.
VideoMLA reports 3.38 seconds latency on B200 in the chunk-wise autoregressive comparison.
What the paper found
VideoMLA, from Virginia Tech, is the first Multi-Head Latent Attention system applied to autoregressive video diffusion, built on the Wan2.1-T2V-1.3B backbone. Instead of caching dense per-head keys and values, it stores a shared low-rank content latent plus a head-shared decoupled 3D-RoPE positional key, shrinking per-token cache state from 3,072 scalars to 224, a 92.7% reduction, or 13.7× smaller memory per cached layer. The paper’s key technical finding is that this works even though video attention is not intrinsically low-rank: on Wan2.1-T2V-1.3B, the 99%-energy effective rank of the dense [WK; WV] operator exceeds 1,300 in every layer, so direct spectral compression would fail. Instead, the MLA bottleneck itself defines the rank budget, and training saturates that budget from initialization, whether starting from SVD or random weights, with the learned operator tracking roughly 0.98dc effective rank across depth. On VBench, VideoMLA matches short-horizon streaming baselines and achieves the best long-horizon score, reaching 0.859 overall at 60 seconds, while improving throughput to 23.96 FPS and cutting latency to 3.38 seconds on a single B200. The work shows that compressing the KV layout, rather than changing the rollout window, is a complementary path to minute-scale video generation.
Original abstract
Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the window or how their positions are encoded. The per-head KV layout itself, a dominant contributor to streaming memory and latency, has been mostly left unchanged. In this paper, we present the first study of Multi-Head Latent Attention (MLA) in video diffusion. VideoMLA replaces per-head keys and values with a shared low-rank content latent and a shared decoupled 3D-RoPE positional key, reducing per-token KV memory by 92.7% at every cached layer. We further investigate why MLA succeeds in video diffusion even though the spectral assumption often used to motivate it in language models does not hold: pretrained video attention is not low-rank, with 99%-energy effective rank far above any practical latent dimension. VideoMLA retains quality at compression ratios where direct spectral approximation would predict large reconstruction error. We show that the MLA bottleneck, rather than the pretrained spectrum, determines the effective rank: both spectral and random initialization occupy nearly the full rank budget from initialization, and training preserves this budget while adapting within it. On VBench, VideoMLA matches short-horizon streaming video diffusion baselines, achieves the best overall score at long horizons among evaluated methods, and improves throughput by 1.23x on a single B200.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.