NTH

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

AuthorsJiapeng Li, Ping Wei, Wenjuan Han, Song-Chun Zhu, Lifeng Fan

September 3, 2026 2 min read
Watch on YouTube
The one-line take

IntentQA pushes video AI beyond recognizing actions toward explaining why people act, using multimodal context and robustness tests to evaluate genuine intent reasoning.

Key results

4,303
IntentQA videos

Number of videos in the IntentQA dataset.

16,297
IntentQA question-answer pairs

Total annotated question-answer pairs.

64.81%
Full GPT-4 accuracy

Main test-set accuracy of X-CaVIR using GPT-4.

3.98%
Confidence Prompting gain

Accuracy increase from adding the InstructGPT Confidence Prompting pipeline.

3.75%
GPT-4 contrast decline

Average Contrast Performance Decline for the full GPT-4 system.

What the paper found

IntentQA reframes VideoQA as inference over latent human goals rather than recognition of visible facts, covering Causal Why, Causal How, Temporal Previous, and Temporal Next questions. Its dataset contains 4,303 videos and 16,297 question-answer pairs derived from NExT-QA, with 624 distinct actions. The proposed X-CaVIR framework combines three cognitive contexts: Situational Context extracted by a Video Query Language module over dynamic region graphs, Contrastive Context learned with WUPS-based positive and negative triplets, and Commonsense Context supplied by a large language model. To test genuine reasoning, the benchmark adds five LLM-generated contrast sets that perturb verbs, nouns, and gender, and introduces Contrast Performance Decline to measure robustness against superficial pattern matching. X-CaVIR uses Confidence Prompting to place VideoQA answer likelihoods inside an LLM prompt alongside dense video captions generated by Qwen 7B, producing both an answer and an explicit reasoning trace. On the main test set, the full InstructGPT system reaches 60.73% accuracy, a 3.98% increase over its preceding visual model, while replacing InstructGPT with ChatGPT gives 59.28% and GPT-4 gives 64.81%. GPT-4 also limits average contrast-set decline to 3.75%, outperforming traditional visual-language baselines, although human accuracy remains 78.49%. The results show that combining visual evidence, contrastive learning, and commonsense reasoning improves intent understanding and interpretability, but subtle social motivations and missing visual details still challenge current LLM-based systems.

Original abstract

Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR (eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the opacity of traditional black-box models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis