MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism
AuthorsCong Chen, Guo Gan, Kaixiang Ji, ChaoYang Zhang, Zhen Yang, Guangming Yao, Hao Chen, Jingdong Chen, Yi Yuan, Chunhua Shen
Resources
MemDreamer lets vision-language models understand hours-long videos by storing them in a hierarchical graph memory and retrieving only the most relevant pieces during reasoning.
Key results
MemDreamer with Gemini-3.1-Pro
Absolute LVBench gain over Gemini-3.1-Pro end-to-end
Distance from human expert on LVBench
LVBench gain from 63.6 to 84.8
Lower end of reasoning context size during retrieval
Upper end of end-to-end input size on long videos
What the paper found
MemDreamer, developed by an Ant Group and Zhejiang University team, reframes long-video understanding as a decoupled perception-and-reasoning system instead of end-to-end token ingestion. A streaming perception model builds a purely textual Hierarchical Graph Memory with three levels—Video Root, Super Events, and Macro Events—plus leaf subgraphs whose entities, micro-events, and edges encode spatial, temporal, and causal structure. At inference, a separate reasoning model uses an agentic Observation-Reason-Action loop with seven tools, including hierarchical navigation, semantic search, time filtering, and graph traversal, so it can retrieve evidence without rereading raw video. On LVBench, the full Gemini-3.1-Pro configuration reaches 90.7, beating the strongest native baseline by 12.5 points and leaving only a 3.7-point gap to human experts; across the four benchmarks it reports 86.3 on LongVideoBench, 92.1 on Video-MME, and 88.2 on EgoSchema. The active reasoning context is only 5.9K to 6.3K tokens, versus 240K to 784K for end-to-end models, and the paper reports a 21.2-point jump for Qwen3-VL, from 63.6 to 84.8, when wrapped in MemDreamer. Ablations show that Hierarchical-Graph memory outperforms Flat-Chunk by 13.3 points, and the full tool suite outperforms vanilla embedding similarity by 20.2 points.
Original abstract
Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduce MemDreamer to decouple perception and reasoning, shifting long-video understanding into an agentic exploration process. As a plug-and-play framework, it incrementally streams videos to construct a Hierarchical Graph Memory, a top-down three-tier architecture for semantic abstraction, anchored by a foundational graph capturing spatiotemporal and causal relations. During inference, the reasoning model employs agentic tool-augmented retrieval, navigating hierarchies, searching nodes, and traversing logical edges via an Observation-Reason-Action loop. Experiments show MemDreamer achieves SOTA results across four mainstream benchmarks, narrowing the gap with human experts to only 3.7 points. It constrains the reasoning context window to merely 2% of full-context ingestion while delivering a 12.5 point absolute accuracy gain. Furthermore, statistical analysis uncovers a strong positive linear correlation between an VLM's performance on logic reasoning and long-video understanding benchmarks, establishing agentic capability scaling as a new paradigm for multimodal comprehension.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.