Linear Scaling Video VLMs for Long Video Understanding
AuthorsCristobal Eyzaguirre, Jiajun Wu, Juan Carlos Niebles
Resources
This paper makes long video understanding cheaper by replacing quadratic video attention with a stateful linear-time cache that preserves much of the accuracy of full self-attention.
Key results
Mean fraction of historical attention mass captured by top-4096 tokens for InternVL3-1B/2B/8B on 16 long videos.
Weighted candidate-pool recall for the next oracle state at B = 256 on the same attention-analysis set, reported for InternVL3-8B.
StateKV result on VideoMME for InternVL3-8B with B = 4096.
Full self-attention baseline on VideoMME for InternVL3-8B.
ReKV baseline on VideoMME for InternVL3-8B with R = 16.
Approximate horizon where InternVL3-8B with the largest StateKV cache becomes cheaper than InternVL3-1B Full SA.
What the paper found
StateKV, from Stanford University, is an inference-time method for pretrained video vision-language models that linearizes long-video prefill without fine-tuning or architecture changes. Instead of letting every new frame attend to the entire past, it maintains two KV caches per layer: a fixed-capacity importance-based temporal state for cross-frame context and a detailed full cache for final decoding, which reduces video-prefill complexity from quadratic to linear in frame count. The paper’s mechanistic analysis on 16 long videos from the VideoMME training split shows that historical attention is highly concentrated: the top-4096 historical tokens capture 0.93 of historical attention mass for InternVL3-1B/2B/8B, and the incremental candidate pool recovers the next oracle state with weighted recall of 0.96, 0.97, and 0.96 at B = 256. On three benchmarks—VideoMME, MLVU, and OVOBench—StateKV with B = 4096 stays close to full self-attention and consistently outperforms ReKV’s 16-frame sliding window; for example, on InternVL3-8B it reaches 62.52% on VideoMME versus 64.19% for full self-attention and 54.56% for ReKV. The compute frontier is especially strong: StateKV-8B at B = 4096 achieves 62.5% VideoMME accuracy at similar FLOPs to Full SA-1B, and the authors report that the larger StateKV model becomes compute-favorable beyond about 1800 frames.
Original abstract
Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal self-attention, causing compute and latency to grow quadratically with the number of frames. Existing efficiency methods improve scalability but often lose accuracy relative to full self-attention, for example through aggressive frame/token dropping or coarse attention approximations. We introduce StateKV, an inference-time method that adapts pretrained long-video VLMs to linear-time video prefill by carrying cross-frame context in a fixed-capacity, importance-based recurrent state, paired with a second full per-frame cache used for decoding. Across three long-video benchmarks and seven models spanning three families and multiple scales, StateKV remains close to full self-attention and consistently outperforms dominant sliding-window / recency-based streaming approximations, without fine-tuning or architectural changes. StateKV also reduces video-prefill cost measured FLOPs, enabling stronger accuracy at a fixed compute budget by running larger models. These results suggest a practical step toward scalable long-video understanding.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.