EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
AuthorsWeitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
Resources
EgoMonth tests whether AI can genuinely remember a person’s daily visual experiences over months rather than merely summarize isolated clips.
Key results
Over 300 hours of first-person video
Final participant count
Multiple-choice questions
Tasks across three cognitive levels
Best model accuracy on EgoMonth
Human baseline accuracy
What the paper found
EgoMonth introduces the first month-level egocentric video benchmark designed to test whether multimodal large language models can maintain faithful spatiotemporal memory across real daily life rather than isolated web clips. The dataset contains over 300 hours of first-person recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice questions. Its 14 tasks are organized into Schema Consolidation, Episodic Indexing, and Cascading Reasoning, covering habits, event timing, object locations, counting, routes, cross-view spatial relations, and directional judgment. Privacy protection combines Grounding DINO 1.5 for sensitive-region detection with SAM 2 for segmentation and masking. Across open and closed models, including Qwen2.5-VL and Google’s Gemini 2.5 Pro, Gemini 2.5 Pro achieves the top macro-average accuracy at 71.8%, while the human baseline reaches 94.2%. Qwen2.5-VL, the strongest open-source model, scores 58.0%. Performance declines as reasoning requires more temporally distant and spatially structured evidence, with several models approaching the 25% four-option chance level on route and cross-view tasks. Results from evaluations using NVIDIA RTX 4090 systems indicate that simply increasing frame density or parameter count does not solve memory failure: models need event-level indexing, selective retrieval, persistent object-state tracking, and structured spatial representations. EgoMonth therefore frames current MLLMs as lossy summarizers rather than reliable long-term memorizers.
Original abstract
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.