Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
AuthorsXiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou
Resources
Xiaomi-Robotics-1 scales robot learning with over 100K hours of language-annotated real-world data to produce a stronger and more adaptable general-purpose manipulation policy.
Key results
Hours of real-world manipulation trajectories used for pre-training.
Approximate hours of robot and instruction-labeled trajectories used for alignment.
Average success rate, compared with the previous best of 46.6%.
Average simulation score, compared with the prior state of the art at 13.07.
Average success across four new tasks with less than 10 hours per task.
Model-scale variant reaching 79% out-of-the-box success after post-training.
What the paper found
Xiaomi Robotics introduces Xiaomi-Robotics-1, a scalable vision-language-action model built around a Mixture-of-Transformers architecture that couples Qwen3-VL with a diffusion transformer using flow matching. Its two-stage recipe first pre-trains on 100k hours of real-world manipulation trajectories captured with Universal Manipulation Interface devices, while Qwen3.5-27B automatically labels trajectory clips with language describing scene-state transitions; post-training then uses about 10k hours of cross-embodiment data to align UMI skills with robot embodiments and imperative instructions. Scaling experiments show that both larger datasets and larger models improve action prediction and transfer to unseen real-world environments: the 5B model’s out-of-the-box success rate rises from 26% without action pre-training to 75% with the full pre-training scale, while model scaling increases success from 61% for 2B to 79% for 10B. For efficient adaptation, Xiaomi-Robotics-1 reaches 75% average success across four dexterous tasks with less than 10 hours per task, versus 40% for π0.5. In simulation, it achieves 74.5% on RoboCasa, 57.4% on RoboCasa365 versus the previous best of 46.6%, and an average RoboDojo score of 20.07 versus 13.07. The results position Xiaomi’s model as a robot foundation policy whose large-scale UMI pre-training improves both zero-shot generalization and fine-tuning efficiency.
Original abstract
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.