NTH

Video-Index: A Curated Meta-Benchmark for Video Understanding

AuthorsEnxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu

Affiliations[

October 6, 2026 2 min read
Watch on YouTube
The one-line take

Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.

Key results

115
Audited benchmarks

Number of video benchmarks evaluated with the five-level attack pyramid.

96%
Shuffled-frame retention

Median fraction of full-video accuracy retained after frame shuffling on 51 benchmarks.

63
Near-duplicate-heavy benchmarks

Benchmarks where near-duplicate questions make up at least half of the items.

840
Video-Index items

Verified hard items released across four capability groups.

56.8%
Claude Opus 5 fixed-input accuracy

Performance on Video-Index under the fixed-input protocol.

19.8%
Tool-use gain

Accuracy improvement from tools on matched items.

What the paper found

Video-Index argues that many video-understanding scores measure shortcuts rather than video comprehension. Its attack pyramid tests five increasingly informed adversaries: answer options, question text, previously evaluated items, a single frame or captions, and shuffled or truncated video. Auditing 115 multiple-choice benchmarks found that 35 can be solved nearly at full-video accuracy without seeing a frame, while on 51 benchmarks shuffled frames retain a median 96% of the original accuracy; near-duplicate questions constitute at least half of the items in 63 benchmarks. The authors screen 505,518 question–answer pairs with Qwen3-VL-8B and Qwen3-VL-2B, then use deterministic, coverage-aware selection, agent-generated specifications, Claude Opus 5 verification, and a red-team gate to release Video-Index: 840 verified hard items, divided into four capability groups with 210 items each and drawn from 76 sources. Under fixed inputs, Anthropic’s Claude Opus 5 reaches 56.8%, compared with 19.0% for the strongest open model, Molmo2-8B, while tool use raises matched-model performance by 19.8%. Agentic systems perform substantially better: OpenAI’s GPT-6-Astra reaches 79.3%, ahead of Claude Fable 5.1 at 71.5%, but performance still degrades on long videos, showing that evidence selection and efficient temporal inspection remain unresolved.

Original abstract

A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: https://www.enxinsong.com/blog/video-index/ GitHub: https://github.com/Espere-1119-Song/Video-Index Hugging Face: https://huggingface.co/datasets/Video-Index/Video-Index

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis
03Benchmark

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan

GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.

Read analysis