NTH

Evidence-Backed Video Question Answering

AuthorsShijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles

July 18, 2026 2 min read
Watch on YouTube
The one-line take

E-VQA teaches video-language models to answer questions while showing exactly when and where in the video the evidence appears.

Key results

160k
ST-Evidence-Instruct scale

Instruction-tuning dataset of aligned QA, temporal evidence, and spatial masklet triplets.

6 FPS
Mask annotation rate

Frame rate used for dense tracked spatio-temporal mask annotations.

27.2
Ours-7B temporal gain

Point improvement in t-mean over the size-matched UniPixel-7B baseline.

13.8
Ours-7B spatial gain

Point improvement in J&F over the size-matched UniPixel-7B baseline.

82.82
Gemini-2.5-Pro MCQ QA accuracy

QA accuracy on ST-Evidence-MCQ.

44.0
Gemini-2.5-Pro generative spatial score

J&F score for spatial evidence on ST-Evidence-Gen.

What the paper found

Researchers from Salesforce and Brown University introduce Evidence-Backed Video Question Answering, or E-VQA, which requires a Video LLM to output not only an answer but also supporting time segments and dense, tracked pixel masks called masklets. Their human-verified ST-Evidence benchmark contains 1,298 questions for multiple-choice grounding and 2,706 questions for generative grounding, with masks annotated at 6 FPS. It draws a sharp distinction between knowing an answer and proving it visually: OpenAI o3 and Gemini-2.5-Pro can achieve strong QA, while spatial grounding remains inconsistent, and many open-source systems perform near the 25.00 random baseline in mask selection. The authors also release ST-Evidence-Instruct, a 160k-scale instruction-tuning dataset built through automated pipelines combining Qwen3-VL, Gemini-2.5-Pro, and SAM-3. Fine-tuning UniPixel-based models with LoRA produces major grounding improvements over size-matched baselines: the 7B model gains 27.2 t-mean points for temporal evidence and 13.8 J&F points for spatial evidence, while QA accuracy rises only modestly. On the multiple-choice benchmark, Gemini-2.5-Pro reaches 82.82 QA accuracy, but its 44.0 J&F generative spatial score shows that semantic competence does not guarantee dense visual justification. The central conclusion is that scaling alone cannot close this perception-reasoning gap; aligned evidence data, integrated objectives, and architectures that jointly reason over time and pixels are necessary for explainable video understanding.

Original abstract

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. To support this, we introduce ST-Evidence, the first human-verified benchmark for both discriminative and generative pixel-level grounding. Evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and true visual perception that scaling alone fails to bridge. To address this, we develop scalable, automated generation pipelines to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding. Fine-tuning grounded Video LLMs on this data yields substantial gains over the corresponding size-matched UniPixel baselines (e.g., +27.2 t-mean and +13.8 J&F on a 7B model), establishing a robust baseline for explainable, evidence-backed video understanding. Code and data are available at https://github.com/SalesforceAIResearch/EVQA.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis