NTH

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

AuthorsHaoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen

July 4, 2026 2 min read
Watch on YouTube
The one-line take

Qwen-RobotManip shows that aligning diverse robot and human demonstration data at scale can unlock stronger generalization for vision-language-action robots across many platforms and real-world settings.

Key results

38100
pretraining corpus hours

approximate manipulation pretraining corpus built from open-source robot and egocentric human data

28M
vision-language mixture size

curated vision-language co-training mixture

72.2%
RoboTwin-IF accuracy

Qwen-RobotManip instruction-following benchmark result

45.6%
EBench overall score

OOD mobile manipulation benchmark success rate

35.9%
RoboCasa365 total success

OOD kitchen manipulation benchmark overall success rate

23.9%
RoboTwin-XE zero-shot total

cross-embodiment transfer success across ARX-X5, UR5-WSG, and Franka Panda

What the paper found

Qwen Team’s Qwen-RobotManip technical report argues that robotic manipulation foundation models only scale when heterogeneous data are first aligned, then expanded. Built on Qwen-VL and Qwen3.5-4B, the system couples a Vision-Language-Action backbone with a flow-matching Diffusion Transformer and unifies manipulation under a canonical 80-dimensional state-action space, camera-frame delta end-effector poses, and in-context policy adaptation. The training corpus is assembled entirely from open-source robot datasets and egocentric human videos, then expanded through a human-to-robot synthesis pipeline across 15 platforms into roughly 38,100 hours of pretraining data, alongside a 28M-point vision-language mixture. On OOD benchmarks chosen to expose genuine generalization rather than memorization, Qwen-RobotManip outperforms π0.5 across LIBERO-Plus, RoboTwin-Clean2Rand, RoboCasa365, EBench, RoboTwin-IF, and RoboTwin-XE, with especially large gains in instruction following and cross-embodiment transfer. It reaches 72.2% on RoboTwin-IF versus 49.6% for π0.5, 45.6% on EBench versus 27.1%, 35.9% on RoboCasa365 versus 16.9%, and 23.9% zero-shot on RoboTwin-XE versus 7.5%. In real-world evaluation, the model also ranks 1st on RoboChallenge Table30-v1 with a 20% relative improvement and shows strong robustness on AgileX ALOHA, Franka, UR, and ARX, validating the paper’s central claim that alignment is what makes large-scale manipulation pretraining transferable.

Original abstract

Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $π$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis