NTH

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

AuthorsQiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Bai, Jingren Zhou, Jiazhao Zhang, Haoqi Yuan, Gengze Zhou, Hang Yin, Ye Wang, Yiyang Huang, Zixing Lei, Wujian Peng, Delin Chen, Yingming Zheng, Jingyang Fan, Xianwei Zhuang, Xin Zhou, Haoyang Li, Anzhe Chen, Tong Zhang, Xuejing Liu, Yuchong Sun, Ruizhe Chen, Zhaohai Li, Chenxu Lü, Zhibo Yang, Tao Yu, Xionghui Chen

June 3, 2026 3 min read
Watch on YouTube
The one-line take

Qwen-VLA is a single robot brain that can see, reason, and act across different tasks and robot bodies by training one vision-language model to handle navigation, manipulation, and trajectory prediction together.

Key results

4B
Qwen-VLA backbone

The model is built on the Qwen3.5-4B multimodal backbone.

1.15B
Action decoder size

The DiT-style flow-matching action expert contains approximately 1.15B parameters.

6.0%
Human egocentric data share

Human egocentric trajectories make up 6.0% of the pretraining mixture.

7.5%
Navigation data share

Navigation trajectories make up 7.5% of the pretraining mixture.

359,848
Synthetic VLA trajectories

The synthetic simulation pipeline produces 359,848 full successful VLA trajectories including subtask segments.

7.2M
Language-only simulated trajectories

The language-only action dataset contains roughly 7.2M simulated trajectories for text-to-action pretraining.

What the paper found

Qwen-VLA, from the Qwen Team at Qwen, presents a single vision-language-action foundation model that unifies manipulation, vision-language navigation, egocentric human motion, and trajectory prediction across robot embodiments. Built on the Qwen3.5-4B multimodal backbone, it adds a 1.15B-parameter DiT-style flow-matching action decoder that generates continuous action chunks with a few Euler steps, while embodiment-aware prompts specify robot platform, arm configuration, control frequency, and horizon so one model can handle heterogeneous control conventions without separate heads. The core novelty is a staged training recipe: text-to-action pretraining teaches the decoder to reconstruct full trajectories from language alone, continued pretraining grounds it with images, supervised fine-tuning aligns downstream tasks, and PPO-based reinforcement learning on sparse success rewards further improves closed-loop control. Training uses a large mixed corpus dominated by real and simulated robot manipulation trajectories, plus 6.0% egocentric human data, 7.5% navigation data, and curated vision-language supervision; the authors also introduce 359,848 synthetic VLA trajectories and 7.2 million language-only simulated trajectories for action prior learning. On benchmarks, Qwen-VLA-Instruct reaches 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% oracle success on R2R, 59.6% success on RxR, 76.9% average out-of-distribution success on real ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation, often beating specialist policies such as π0.5, GR00T N1.6, and StarVLA. The paper’s main claim is that embodied intelligence can be scaled like multimodal pretraining if actions are treated as a shared generative prediction problem rather than isolated robot-specific policies.

Original abstract

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and auxiliary vision-language data. To support multiple robot platforms, we introduce embodiment-aware prompt conditioning, where robot-specific textual descriptions specify the current embodiment and control convention. We further cast manipulation, navigation, and trajectory prediction into a unified action-and-trajectory prediction framework, enabling transferable visual grounding, spatial reasoning, and continuous action generation across robot morphologies, task families, and environments. Experiments on manipulation, navigation, and trajectory-centric benchmarks show consistent multi-task performance and out-of-distribution generalization under variations in scene layout, background, lighting, object configuration, and robot embodiment. Qwen-VLA-Instruct achieves 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% OSR on R2R, 59.6% SR on RxR, 76.9% average OOD success in real-world ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis