M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
AuthorsJie Huang, Ruixun Liu, Sirui Sun, Xinyi Yang, Yin Li, Yixin Zhu, Yiwu Zhong
Resources
M^3Eval is a new benchmark that tests whether multimodal models can actually remember what they see and hear in long videos, revealing surprising weaknesses in how they store and separate information.
Key results
The benchmark comprises 2,403 questions across all memory tasks.
The benchmark uses 451 source videos for the non-N-Back tasks.
The source videos span approximately 403 hours.
What the paper found
M^3Eval, developed by Jie Huang, Ruixun Liu, Sirui Sun, Xinyi Yang, Yin Li, Yixin Zhu, and Yiwu Zhong at Peking University and the University of Wisconsin-Madison, is the first benchmark to isolate multi-modal memory in video understanding rather than conflating it with perception or reasoning. Grounded in cognitive psychology, it defines four memory dimensions: divided attention for concurrent streams, interference for retroactive versus proactive distraction, interleaved events for temporal organization, and N-Back for symbolic grounding and memory capacity. The benchmark spans 2,403 questions over 451 videos and about 403 hours, drawn from HourVideo, Video-MME, LVBench, InfiniBench, and CrossVid, with questions generated by Qwen3.5-27B and manually verified. Across proprietary and open models including Google DeepMind’s Gemini-3.1-Pro-Preview, OpenAI’s GPT-5.4, Qwen3.5, Qwen3-VL, InternVL3.5, VideoLucy, and M3-Agent, the results show substantial gaps from human performance: models often fall near chance on split-screen source identification, exhibit high intrusion rates under semantically similar interference, fail to reconstruct interleaved temporal order, and lag far behind humans on N-Back. A notable finding is that repeating target or interfering videos can improve accuracy, while temporal source grounding is consistently weaker than spatial grounding, revealing a core limitation in current transformer-based memory mechanisms.
Original abstract
As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmarks, existing works primarily focus on perception and reasoning, without systematically evaluating memory: what models retain, how faithfully information is preserved, and how robust memory remains under interference. To address this gap, we introduce M$^3$Eval, the first comprehensive evaluation framework and benchmark for probing different memory dimensions in multi-modal models. Grounded in cognitive psychology, our design features carefully constructed tasks that isolate key aspects of memory. Leveraging M$^3$Eval, we conduct extensive experiments across representative multi-modal models, revealing consistent weaknesses and distinctive behaviors. We find that models struggle to maintain disentangled representations when processing parallel video streams, exhibit interference patterns differing substantially from those observed in human memory, ground memory sources more reliably in the spatial domain than the temporal domain, and demonstrate limited symbolic memory. Collectively, our benchmark provides a valuable resource for future research, while our findings highlight memory as a fundamental yet underexplored capability and offer insights for designing more effective memory mechanisms in multi-modal models. Our code and dataset are available at https://pku-value-lab.github.io/m3eval-homepage.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.