NTH

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

AuthorsHidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay, Hoda Eldardiry, Pinar Yanardag

June 1, 2026 2 min read
Watch on YouTube
The one-line take

VideoMLA makes minute-scale autoregressive video diffusion much cheaper in memory by compressing the KV cache with a low-rank latent representation, enabling longer video rollouts with better throughput and strong quality.

Key results

92.7%
KV cache reduction

VideoMLA reduces per-token cached KV state from 3072 dense scalars to 224 scalars per layer.

224
Per-token cache size

Default VideoMLA setting stores 224 scalars per cached token per layer using a 192-dim content latent plus 32 RoPE channels.

3072
Dense KV cache size

Wan2.1-T2V-1.3B dense per-token, per-layer KV cache stores 2 × 12 × 128 = 3072 scalars.

0.859
VBench 60s overall

VideoMLA achieves the best overall VBench score at 60 seconds among evaluated methods.

23.96
Throughput

VideoMLA reaches 23.96 FPS on B200 in the chunk-wise autoregressive comparison.

3.38
Latency

VideoMLA reports 3.38 seconds latency on B200 in the chunk-wise autoregressive comparison.

What the paper found

VideoMLA, from Virginia Tech, is the first Multi-Head Latent Attention system applied to autoregressive video diffusion, built on the Wan2.1-T2V-1.3B backbone. Instead of caching dense per-head keys and values, it stores a shared low-rank content latent plus a head-shared decoupled 3D-RoPE positional key, shrinking per-token cache state from 3,072 scalars to 224, a 92.7% reduction, or 13.7× smaller memory per cached layer. The paper’s key technical finding is that this works even though video attention is not intrinsically low-rank: on Wan2.1-T2V-1.3B, the 99%-energy effective rank of the dense [WK; WV] operator exceeds 1,300 in every layer, so direct spectral compression would fail. Instead, the MLA bottleneck itself defines the rank budget, and training saturates that budget from initialization, whether starting from SVD or random weights, with the learned operator tracking roughly 0.98dc effective rank across depth. On VBench, VideoMLA matches short-horizon streaming baselines and achieves the best long-horizon score, reaching 0.859 overall at 60 seconds, while improving throughput to 23.96 FPS and cutting latency to 3.38 seconds on a single B200. The work shows that compressing the KV layout, rather than changing the rollout window, is a complementary path to minute-scale video generation.

Original abstract

Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the window or how their positions are encoded. The per-head KV layout itself, a dominant contributor to streaming memory and latency, has been mostly left unchanged. In this paper, we present the first study of Multi-Head Latent Attention (MLA) in video diffusion. VideoMLA replaces per-head keys and values with a shared low-rank content latent and a shared decoupled 3D-RoPE positional key, reducing per-token KV memory by 92.7% at every cached layer. We further investigate why MLA succeeds in video diffusion even though the spectral assumption often used to motivate it in language models does not hold: pretrained video attention is not low-rank, with 99%-energy effective rank far above any practical latent dimension. VideoMLA retains quality at compression ratios where direct spectral approximation would predict large reconstruction error. We show that the MLA bottleneck, rather than the pretrained spectrum, determines the effective rank: both spectral and random initialization occupy nearly the full rank budget from initialization, and training preserves this budget while adapting within it. On VBench, VideoMLA matches short-horizon streaming video diffusion baselines, achieves the best overall score at long horizons among evaluated methods, and improves throughput by 1.23x on a single B200.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis