MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs
AuthorsDhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas
Resources
MotionBlind shows that many Video-LLMs can recognize what is happening in a clip while still failing to understand how fast or in which direction it happens.
Key results
Contrastive benchmark instances spanning speed, magnitude, and direction.
Binary items forming the benchmark’s strict 2×2 evaluations.
Chance level when all four answers in an instance must be correct.
Eagle2.5-8B’s best MotionBlind result.
Only frontier model reported to clear the benchmark overall.
Mean performance of human annotators on MotionBlind.
What the paper found
MotionBlind tests whether Video-LLMs genuinely perceive physical motion rather than infer answers from appearance or language priors. Its contrastive benchmark contains 60 instances, 82 self-recorded clips, and 240 yes/no items covering speed, magnitude, and direction; each instance uses two near-identical videos and requires all four paired answers to be correct, making chance performance 6.25%. Across six open models, including Eagle2.5-8B, Qwen3-VL-4B, Motion-o, Gemma-4-12B-it, Molmo2-8B, and Inkling-Small, the best MotionBlind score is only 11.7% IAcc, despite much stronger results on conventional Video-MME. Removing the video produces exactly 0% IAcc, while shuffling or reversing frames reduces performance to chance, confirming that ordered visual evidence is necessary but not sufficient. Increasing the frame budget from 1 to 24 or replacing uniform sampling with HORNet, a GRPO-trained selector, or Frame2Clip does not solve the deficit; models plateau near 12%. The main exception is Google’s Gemini 3.1 Pro, which reaches 60.0% IAcc, although it falls to 14% on speed, showing that rate estimation remains difficult; OpenAI’s GPT-5.6 Luna reaches only 15.0%. Human annotators achieve 91.3% IAcc, exposing a large perceptual gap. The study argues that Video-LLMs should not yet be trusted as motion-aware supervision, reward, or evaluation components for physical-AI world models.
Original abstract
Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, which way it travels, how hard it is pushed. We show they cannot. A Video-LLM can watch two clips of the same person in the same room, name every object in both, and still fail to say which clip moves faster. We introduce MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion(speed, magnitude, and direction), the variables a world model must predict. Each instance is a pair of near-identical clips that differ only in motion. Each clip carries two complementary yes/no questions, giving four items per instance, and a model earns credit only if all four are correct. We report Instance Accuracy(IAcc), which has a 6.25% chance floor. Single-frame, appearance, and language-only shortcuts all collapse to it. MotionBlind complements the recent TimeBlind benchmark. We run a controlled study of six open and two frontier Video-LLMs, varying whether the video is present, whether frames are shown in the correct temporal order, and how frames are sampled (1 to 24 frames, four selection strategies). Open models sit near the 6.25% floor, and scale does not help. Removing the video drops every model to zero IAcc, and shuffling frames collapses IAcc to chance, so the task genuinely needs video in order. Neither more frames nor smarter frame selection closes the gap, because these change which frames are seen, not whether motion is read. Only Gemini3.1 Pro clears the benchmark overall, and even it fails on speed. A frontend that cannot tell two speeds of the same action apart is not yet a trustworthy source of supervision, reward, or evaluation for a world model.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.