VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
AuthorsYunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang
VIABench tests whether multimodal AI can genuinely assist blind users in real-world video scenarios, rather than merely answer questions about static images.
Key results
Number of first-person videos in the benchmark.
Number of manually curated annotations.
Total benchmark footage in hours.
Best overall score reported across VIABench evaluation tasks.
Speedup over repeated prompting at identical detection accuracy.
What the paper found
VIABench, from Nanjing University and Shanghai AI Laboratory, is a video benchmark designed around real visual-assistance needs reported by blind users rather than conventional curated video understanding. It contains 761 first-person videos, 14526 manually curated annotations, and 46.9 hours of footage, including long clips averaging 222 seconds. The benchmark unifies three core tasks: Proactive Reminder, which requires timely, autonomous alerts for hazards and navigation events; Visual Question Answering, which answers questions using only video observed before the query; and Vision-Guided Interaction, which provides iterative, closed-loop instructions for tasks such as locating elevator buttons. Its key evaluation method, Token-Level Prompt Activation Decoding, or TPAD, converts offline multimodal large language models into frame-wise alert detectors without fine-tuning by extracting token-level alert probabilities in a single forward pass; a controlled test reports a 4.5x speedup over repeated prompting at identical detection accuracy. Evaluations cover open models including InternVL3.5, Qwen2.5VL, LLaVA, and MiniCPM, streaming systems such as LiveCC and VideoLLM-Online, and proprietary models including OpenAI’s GPT-5 and GPT-4o and Google’s Gemini-2.5 Pro. Results expose a major deployment gap: GPT-5, the strongest evaluated model, achieves only 28.8 overall, while proactive alerting and multi-turn guidance remain much weaker than ordinary VQA. The authors conclude that assistive models need blind-centered training, stronger temporal anticipation, lower latency, and better robustness to blurry, occluded, rotated, or misexposed egocentric video.
Original abstract
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.