NTH

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

AuthorsHongyu Qu, Jianzhe Gao, Xiaobin Hu, Shaohuan Yang, Xinlei Yu, Rui Yan, Wenguan Wang, Xiangbo Shu, Shuicheng Yan

July 13, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches robot vision-language-action models to remember past experience in a shared latent space, helping them handle longer and more complex manipulation tasks.

Key results

73.9%
SimplerEnv-Bridge Avg. Success

LaMem-VLA average success on SimplerEnv-Bridge

97.6%
LIBERO Avg. Success

LaMem-VLA average success across five LIBERO suites

97.0%
LIBERO-90 Success

LaMem-VLA success on the LIBERO-90 suite

7B
Backbone size

Prismatic vision-language model backbone size

300M
Action expert size

Approximate parameter count of the diffusion action expert

57.3%
No dual memory on SimplerEnv

Ablation without both short-term and long-term latent memory

What the paper found

LaMem-VLA introduces a dual latent memory architecture for robotic manipulation that moves historical experience into the native embedding space of vision-language-action models instead of treating memory as external policy-side context. Built on a 7B-parameter Prismatic VLM backbone and a diffusion action expert with approximately 300M parameters, it splits history into a short-term visual vault and a long-term semantic/action-continuity vault, retrieves task-relevant evidence with a context-aware query, compresses it into fixed-length latent tokens, and weaves those tokens directly into the reasoning sequence before action decoding. This design targets the Markovian short-horizon bias that limits long-horizon manipulation. On SimplerEnv-Bridge, LaMem-VLA reaches 73.9% average success, outperforming CogACT by 16.6 points and π0 by 4.7 points. On LIBERO, it achieves 97.6% average success across five suites, beating MemoryVLA by 1.1 points and CogACT by 4.4 points, while reaching 97.0% on LIBERO-90. Ablations show that removing both memory streams drops performance to 57.3% on SimplerEnv and 92.1% on LIBERO-90, confirming that latent-native integration, not just retrieval, is the key gain.

Original abstract

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis