NTH

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

AuthorsCong Chen, Guo Gan, Kaixiang Ji, ChaoYang Zhang, Zhen Yang, Guangming Yao, Hao Chen, Jingdong Chen, Yi Yuan, Chunhua Shen

June 14, 2026 2 min read
Watch on YouTube
The one-line take

MemDreamer lets vision-language models understand hours-long videos by storing them in a hierarchical graph memory and retrieving only the most relevant pieces during reasoning.

Key results

90.7
LVBench score

MemDreamer with Gemini-3.1-Pro

12.5
Improvement over strongest baseline

Absolute LVBench gain over Gemini-3.1-Pro end-to-end

3.7
Human gap

Distance from human expert on LVBench

21.2
Qwen3-VL improvement

LVBench gain from 63.6 to 84.8

5.9K
Active context window

Lower end of reasoning context size during retrieval

784K
Full-context input window

Upper end of end-to-end input size on long videos

What the paper found

MemDreamer, developed by an Ant Group and Zhejiang University team, reframes long-video understanding as a decoupled perception-and-reasoning system instead of end-to-end token ingestion. A streaming perception model builds a purely textual Hierarchical Graph Memory with three levels—Video Root, Super Events, and Macro Events—plus leaf subgraphs whose entities, micro-events, and edges encode spatial, temporal, and causal structure. At inference, a separate reasoning model uses an agentic Observation-Reason-Action loop with seven tools, including hierarchical navigation, semantic search, time filtering, and graph traversal, so it can retrieve evidence without rereading raw video. On LVBench, the full Gemini-3.1-Pro configuration reaches 90.7, beating the strongest native baseline by 12.5 points and leaving only a 3.7-point gap to human experts; across the four benchmarks it reports 86.3 on LongVideoBench, 92.1 on Video-MME, and 88.2 on EgoSchema. The active reasoning context is only 5.9K to 6.3K tokens, versus 240K to 784K for end-to-end models, and the paper reports a 21.2-point jump for Qwen3-VL, from 63.6 to 84.8, when wrapped in MemDreamer. Ablations show that Hierarchical-Graph memory outperforms Flat-Chunk by 13.3 points, and the full tool suite outperforms vanilla embedding similarity by 20.2 points.

Original abstract

Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduce MemDreamer to decouple perception and reasoning, shifting long-video understanding into an agentic exploration process. As a plug-and-play framework, it incrementally streams videos to construct a Hierarchical Graph Memory, a top-down three-tier architecture for semantic abstraction, anchored by a foundational graph capturing spatiotemporal and causal relations. During inference, the reasoning model employs agentic tool-augmented retrieval, navigating hierarchies, searching nodes, and traversing logical edges via an Observation-Reason-Action loop. Experiments show MemDreamer achieves SOTA results across four mainstream benchmarks, narrowing the gap with human experts to only 3.7 points. It constrains the reasoning context window to merely 2% of full-context ingestion while delivering a 12.5 point absolute accuracy gain. Furthermore, statistical analysis uncovers a strong positive linear correlation between an VLM's performance on logic reasoning and long-video understanding benchmarks, establishing agentic capability scaling as a new paradigm for multimodal comprehension.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis