NTH

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

AuthorsPrakhar Khatri

September 12, 2026 3 min read
Watch on YouTube
The one-line take

For long-video AI, choosing the right frames—and reinvesting saved visual tokens—matters far more than simply compressing what the model sees.

Key results

6.9
LongVideoBench long-video selection gain

Eight OMP-selected frames beat sixteen uniformly sampled frames by 6.9 points.

11.81
Cross-benchmark OMP gain maximum

OMP improves over uniform sampling by 5.69 to 11.81 points across the three benchmarks.

0.44
Compression accuracy change

Halving the per-frame spatial budget changes accuracy by at most 0.44 points.

2.24
LongVideoBench reinvestment gain

Sixteen compressed OMP frames improve LongVideoBench by 2.24 points over eight full-resolution frames.

3.04
LVBench reinvestment gain

Sixteen compressed OMP frames improve LVBench by 3.04 points over eight full-resolution frames.

What the paper found

This paper isolates how long-video multimodal language models should spend a fixed visual-token budget by changing one factor at a time: frame selection, spatial compression, or reinvestment into additional timestamps. In a shared harness using LongCLIP scores, Qwen3-VL-8B-Instruct, and the LongVideoBench, Video-MME, and LVBench benchmarks, the unmodified Orthogonal Matching Pursuit, or OMP, selects query-relevant and nonredundant frames and consistently outperforms uniform sampling. On hour-long LongVideoBench videos, eight OMP-selected frames beat sixteen uniformly sampled frames by 6.9 points, while OMP’s advantage over uniform sampling ranges from 5.69 to 11.81 points across the three benchmarks and remains within one point of the purpose-built LDDR selector. Replacing LongCLIP with SigLIP changes most selected frames but preserves the selector ranking, indicating that the result is not specific to one scorer. Spatial compression is nearly free: halving each frame’s visual budget changes accuracy by at most 0.44 points at fixed timestamps. The saved tokens become valuable when reinvested into temporal coverage: sixteen compressed OMP frames improve LongVideoBench by 2.24 points and LVBench by 3.04 points over eight full-resolution frames at equal or lower measured token cost. Tests with InternVL3-2B, InternVL3-8B, and GPT-5-mini show that allocation gains depend on the answerer and benchmark, but support the practical order: select relevant frames first, compress second, and spend the savings on more moments.

Original abstract

Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis