Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
AuthorsPrakhar Khatri
Resources
For long-video AI, choosing the right frames—and reinvesting saved visual tokens—matters far more than simply compressing what the model sees.
Key results
Eight OMP-selected frames beat sixteen uniformly sampled frames by 6.9 points.
OMP improves over uniform sampling by 5.69 to 11.81 points across the three benchmarks.
Halving the per-frame spatial budget changes accuracy by at most 0.44 points.
Sixteen compressed OMP frames improve LongVideoBench by 2.24 points over eight full-resolution frames.
Sixteen compressed OMP frames improve LVBench by 3.04 points over eight full-resolution frames.
What the paper found
This paper isolates how long-video multimodal language models should spend a fixed visual-token budget by changing one factor at a time: frame selection, spatial compression, or reinvestment into additional timestamps. In a shared harness using LongCLIP scores, Qwen3-VL-8B-Instruct, and the LongVideoBench, Video-MME, and LVBench benchmarks, the unmodified Orthogonal Matching Pursuit, or OMP, selects query-relevant and nonredundant frames and consistently outperforms uniform sampling. On hour-long LongVideoBench videos, eight OMP-selected frames beat sixteen uniformly sampled frames by 6.9 points, while OMP’s advantage over uniform sampling ranges from 5.69 to 11.81 points across the three benchmarks and remains within one point of the purpose-built LDDR selector. Replacing LongCLIP with SigLIP changes most selected frames but preserves the selector ranking, indicating that the result is not specific to one scorer. Spatial compression is nearly free: halving each frame’s visual budget changes accuracy by at most 0.44 points at fixed timestamps. The saved tokens become valuable when reinvested into temporal coverage: sixteen compressed OMP frames improve LongVideoBench by 2.24 points and LVBench by 3.04 points over eight full-resolution frames at equal or lower measured token cost. Tests with InternVL3-2B, InternVL3-8B, and GPT-5-mini show that allocation gains depend on the answerer and benchmark, but support the practical order: select relevant frames first, compress second, and spend the savings on more moments.
Original abstract
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.