EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
AuthorsGanlin Yang, Zhangzheng Tu, Yuqiang Yang, Sitong Mao, Junyi Dong, Tianxing Chen, Jiaqi Peng, Jing Xiong, Jiafei Cao, Jifeng Dai, Wengang Zhou, Yao Mu, Tai Wang
Resources
EventVLA helps robots remember the right visual moments at the right time, improving long-horizon manipulation by storing sparse, task-critical evidence instead of all past frames.
Key results
Memory-requiring simulation tasks evaluated in the main claim
Bimanual robot tasks evaluated on ARX ACONE
Strictly non-Markovian benchmark tasks
EventVLA with visual anchors plus KEM
EventVLA with visual anchors only
Average improvement over state-of-the-art memory-augmented VLAs
What the paper found
EventVLA, from the Shanghai AI Laboratory and universities including the University of Science and Technology of China, Shanghai Jiao Tong University, Dalian University of Technology, Tsinghua University, Peking University, and Huawei Technologies, reframes long-horizon robotic manipulation as sparse visual evidence memory rather than dense history buffering. Built on Qwen3-VL-4B-Instruct with an optimized fine-tuning policy, it combines fixed visual anchors with a Keyframe Evidence Memory module that predicts future keyframe probabilities from transformer hidden states and commits only task-critical frames through thresholding, 1D non-maximum suppression, and FIFO-bounded storage. The paper also introduces RoboTwin-MeM, a diagnostic benchmark of 8 strictly non-Markovian simulation tasks with 430 to 1,544 average steps per episode and required intermediate keyframes ranging from 1 to 5, plus 4 real-world bimanual tasks on the ARX ACONE robot. Across 17 simulation tasks and 4 real-world tasks, EventVLA reports an average success-rate gain of +40% over prior memory-augmented VLAs; on RoboTwin-MeM it reaches 75.2% versus 18.0% for visual anchors alone, while on RMBench it achieves 67.8%, establishing state-of-the-art performance for conventional memory-oriented manipulation. The key technical novelty is foresight-driven memory writing: instead of reacting after evidence disappears, the model predicts which future steps will become keyframes and stores those observations before occlusion, enabling robust long-horizon control without a redundant memory bank.
Original abstract
Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse, task-critical visual events. This foresight-driven mechanism empowers the policy to dynamically evaluate the future causal utility of current observations, preserving transient visual evidence before it becomes unobservable. Furthermore, we propose RoboTwin-MeM, a diagnostic benchmark specifically designed to evaluate non-Markovian manipulation tasks with interactive visual evidence. Extensive evaluations show that across 17 memory-requiring simulation tasks and 4 real-world bimanual tasks, EventVLA achieves an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.