InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
AuthorsHaoxiang Ma, Junhao Cai, Xiaoxu Xu, Hao Li, Yuyin Yang, Yang Tian, Jiafei Cao, Hongrui Zhu, Zherui Qiu, Zhaxizhuoma, Yuqiang Yang, Jiaqi Peng, Xueyuan Wei, Yangkun Zhu, Jiahao Jiang, Xing Gao, Hanqing Wang, Feng Yuan, Kailin Li, Xueyue Zhu, Tai Wang, Yan Ding, Jiangmiao Pang, Jia Zeng, Jingjing Zhang, Bowen Zhou, Yao Mu, Chunhua Shen, Weinan Zhang
Resources
This paper introduces a robot control model that keeps a vision-language model’s understanding intact while adding compact future prediction, improving long-horizon manipulation and compositional generalization.
Key results
robot manipulation pretraining corpus
co-training corpus for VLM semantics and grounding
Qwen-3.5 model size
lightweight action and foresight module size
learnable latent queries for future prediction
simulation benchmark success rate (%)
What the paper found
InternVLA-A1.5 from the Physical Intelligence Team at Shanghai AI Laboratory unifies vision-language understanding, latent foresight, and continuous action for robot manipulation by building on a native Qwen-3.5 2B VLM backbone and adding a lightweight 460M unified expert. Instead of learning future frames directly in pixel space, it turns prediction into a latent-querying problem with 50 learnable foresight tokens that are supervised through a frozen WAN2.2-5B video generator, so the policy inherits dynamics priors without paying inference-time generation cost. The model is pretrained on 1.2M robot episodes and about 3M multimodal samples, with stage-1 and stage-2 training runs of 300K and 600K steps, and it uses a single next-token objective for VQA, subtask prediction, and discrete action tokens before switching to flow-matching for continuous action chunks. Across six simulation benchmarks, it reaches 98.9 on LIBERO, 84.8 on LIBERO-Plus, 93.2 on RoboTwin 2.0, 49.5 on EBench, and 80.8 on SimplerEnv, while in real-world tests it posts 75.9 on Sort Tubes, 80.5 on Insert Tubes, 76.4 on Move Tubes, and 76.4 on the 13-step MOF procedure, clearly outperforming π0.5 and Motus on the hardest compositional and long-horizon settings.
Original abstract
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing designs tend to erode the semantics of the pretrained backbone, suffer interference among heterogeneous objectives, and learn future prediction from scratch in pixel space, leaving the dynamics priors of pretrained video generators unexploited. We present InternVLA-A1.5, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation. Future prediction is recast as a latent-querying problem, where a small set of learnable foresight tokens condenses the task-relevant future into a compact latent code under the supervision of a frozen pretrained video generation model, so the policy inherits world-model dynamics priors without ever learning pixel-level generation. The video branch is discarded at inference, keeping real-time control. Pretrained on 1.2M robot episodes and 3M multimodal samples, InternVLA-A1.5 achieves the best overall results on all six simulation benchmarks. In the real world, the preserved semantics deliver the strongest compositional generalization on held-out instruction bindings, and the two designs together sustain long-horizon execution.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.