NTH

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

AuthorsJie Huang, Ruixun Liu, Sirui Sun, Xinyi Yang, Yin Li, Yixin Zhu, Yiwu Zhong

June 8, 2026 2 min read
Watch on YouTube
The one-line take

M^3Eval is a new benchmark that tests whether multimodal models can actually remember what they see and hear in long videos, revealing surprising weaknesses in how they store and separate information.

Key results

2403
benchmark questions

The benchmark comprises 2,403 questions across all memory tasks.

451
benchmark videos

The benchmark uses 451 source videos for the non-N-Back tasks.

403
benchmark hours

The source videos span approximately 403 hours.

What the paper found

M^3Eval, developed by Jie Huang, Ruixun Liu, Sirui Sun, Xinyi Yang, Yin Li, Yixin Zhu, and Yiwu Zhong at Peking University and the University of Wisconsin-Madison, is the first benchmark to isolate multi-modal memory in video understanding rather than conflating it with perception or reasoning. Grounded in cognitive psychology, it defines four memory dimensions: divided attention for concurrent streams, interference for retroactive versus proactive distraction, interleaved events for temporal organization, and N-Back for symbolic grounding and memory capacity. The benchmark spans 2,403 questions over 451 videos and about 403 hours, drawn from HourVideo, Video-MME, LVBench, InfiniBench, and CrossVid, with questions generated by Qwen3.5-27B and manually verified. Across proprietary and open models including Google DeepMind’s Gemini-3.1-Pro-Preview, OpenAI’s GPT-5.4, Qwen3.5, Qwen3-VL, InternVL3.5, VideoLucy, and M3-Agent, the results show substantial gaps from human performance: models often fall near chance on split-screen source identification, exhibit high intrusion rates under semantically similar interference, fail to reconstruct interleaved temporal order, and lag far behind humans on N-Back. A notable finding is that repeating target or interfering videos can improve accuracy, while temporal source grounding is consistently weaker than spatial grounding, revealing a core limitation in current transformer-based memory mechanisms.

Original abstract

As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmarks, existing works primarily focus on perception and reasoning, without systematically evaluating memory: what models retain, how faithfully information is preserved, and how robust memory remains under interference. To address this gap, we introduce M$^3$Eval, the first comprehensive evaluation framework and benchmark for probing different memory dimensions in multi-modal models. Grounded in cognitive psychology, our design features carefully constructed tasks that isolate key aspects of memory. Leveraging M$^3$Eval, we conduct extensive experiments across representative multi-modal models, revealing consistent weaknesses and distinctive behaviors. We find that models struggle to maintain disentangled representations when processing parallel video streams, exhibit interference patterns differing substantially from those observed in human memory, ground memory sources more reliably in the spatial domain than the temporal domain, and demonstrate limited symbolic memory. Collectively, our benchmark provides a valuable resource for future research, while our findings highlight memory as a fundamental yet underexplored capability and offer insights for designing more effective memory mechanisms in multi-modal models. Our code and dataset are available at https://pku-value-lab.github.io/m3eval-homepage.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis