Video-Index: A Curated Meta-Benchmark for Video Understanding
AuthorsEnxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Affiliations[
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.
Key results
Number of video benchmarks evaluated with the five-level attack pyramid.
Median fraction of full-video accuracy retained after frame shuffling on 51 benchmarks.
Benchmarks where near-duplicate questions make up at least half of the items.
Verified hard items released across four capability groups.
Performance on Video-Index under the fixed-input protocol.
Accuracy improvement from tools on matched items.
What the paper found
Video-Index argues that many video-understanding scores measure shortcuts rather than video comprehension. Its attack pyramid tests five increasingly informed adversaries: answer options, question text, previously evaluated items, a single frame or captions, and shuffled or truncated video. Auditing 115 multiple-choice benchmarks found that 35 can be solved nearly at full-video accuracy without seeing a frame, while on 51 benchmarks shuffled frames retain a median 96% of the original accuracy; near-duplicate questions constitute at least half of the items in 63 benchmarks. The authors screen 505,518 question–answer pairs with Qwen3-VL-8B and Qwen3-VL-2B, then use deterministic, coverage-aware selection, agent-generated specifications, Claude Opus 5 verification, and a red-team gate to release Video-Index: 840 verified hard items, divided into four capability groups with 210 items each and drawn from 76 sources. Under fixed inputs, Anthropic’s Claude Opus 5 reaches 56.8%, compared with 19.0% for the strongest open model, Molmo2-8B, while tool use raises matched-model performance by 19.8%. Agentic systems perform substantially better: OpenAI’s GPT-6-Astra reaches 79.3%, ahead of Claude Fable 5.1 at 71.5%, but performance still degrades on long videos, showing that evidence selection and efficient temporal inspection remain unresolved.
Original abstract
A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: https://www.enxinsong.com/blog/video-index/ GitHub: https://github.com/Espere-1119-Song/Video-Index Hugging Face: https://huggingface.co/datasets/Video-Index/Video-Index
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.