Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
AuthorsQiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Bai, Jingren Zhou, Jiazhao Zhang, Haoqi Yuan, Gengze Zhou, Hang Yin, Ye Wang, Yiyang Huang, Zixing Lei, Wujian Peng, Delin Chen, Yingming Zheng, Jingyang Fan, Xianwei Zhuang, Xin Zhou, Haoyang Li, Anzhe Chen, Tong Zhang, Xuejing Liu, Yuchong Sun, Ruizhe Chen, Zhaohai Li, Chenxu Lü, Zhibo Yang, Tao Yu, Xionghui Chen
Resources
Qwen-VLA is a single robot brain that can see, reason, and act across different tasks and robot bodies by training one vision-language model to handle navigation, manipulation, and trajectory prediction together.
Key results
The model is built on the Qwen3.5-4B multimodal backbone.
The DiT-style flow-matching action expert contains approximately 1.15B parameters.
Human egocentric trajectories make up 6.0% of the pretraining mixture.
Navigation trajectories make up 7.5% of the pretraining mixture.
The synthetic simulation pipeline produces 359,848 full successful VLA trajectories including subtask segments.
The language-only action dataset contains roughly 7.2M simulated trajectories for text-to-action pretraining.
What the paper found
Qwen-VLA, from the Qwen Team at Qwen, presents a single vision-language-action foundation model that unifies manipulation, vision-language navigation, egocentric human motion, and trajectory prediction across robot embodiments. Built on the Qwen3.5-4B multimodal backbone, it adds a 1.15B-parameter DiT-style flow-matching action decoder that generates continuous action chunks with a few Euler steps, while embodiment-aware prompts specify robot platform, arm configuration, control frequency, and horizon so one model can handle heterogeneous control conventions without separate heads. The core novelty is a staged training recipe: text-to-action pretraining teaches the decoder to reconstruct full trajectories from language alone, continued pretraining grounds it with images, supervised fine-tuning aligns downstream tasks, and PPO-based reinforcement learning on sparse success rewards further improves closed-loop control. Training uses a large mixed corpus dominated by real and simulated robot manipulation trajectories, plus 6.0% egocentric human data, 7.5% navigation data, and curated vision-language supervision; the authors also introduce 359,848 synthetic VLA trajectories and 7.2 million language-only simulated trajectories for action prior learning. On benchmarks, Qwen-VLA-Instruct reaches 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% oracle success on R2R, 59.6% success on RxR, 76.9% average out-of-distribution success on real ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation, often beating specialist policies such as π0.5, GR00T N1.6, and StarVLA. The paper’s main claim is that embodied intelligence can be scaled like multimodal pretraining if actions are treated as a shared generative prediction problem rather than isolated robot-specific policies.
Original abstract
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and auxiliary vision-language data. To support multiple robot platforms, we introduce embodiment-aware prompt conditioning, where robot-specific textual descriptions specify the current embodiment and control convention. We further cast manipulation, navigation, and trajectory prediction into a unified action-and-trajectory prediction framework, enabling transferable visual grounding, spatial reasoning, and continuous action generation across robot morphologies, task families, and environments. Experiments on manipulation, navigation, and trajectory-centric benchmarks show consistent multi-task performance and out-of-distribution generalization under variations in scene layout, background, lighting, object configuration, and robot embodiment. Qwen-VLA-Instruct achieves 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% OSR on R2R, 59.6% SR on RxR, 76.9% average OOD success in real-world ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.