Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
AuthorsXinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
Resources
Xiaomi-Robotics-U0 uses a large multimodal world foundation model to generate consistent robot-centered scenes and videos while improving real-world manipulation performance.
Key results
Multimodal autoregressive parameters in Xiaomi-Robotics-U0
Samples used for general and embodied one-step generation
Video clips used for sequential embodied generation
Maximum speedup from FlashAR-style decoding over standard autoregressive generation
π0.5 progress after augmentation, up from 36.9%
Overall score, ranking first among more than 100 submissions
What the paper found
Xiaomi Robotics introduces Xiaomi-Robotics-U0, a 38B-parameter multimodal autoregressive world model that unifies text-to-image generation, image editing, multi-view embodied scene synthesis, structured embodied transfer, and robot video prediction. Initialized from EMU3.5 and trained with a shared next-token objective, it preserves general visual knowledge while adding robot-specific geometry, camera consistency, embodiment constraints, and interaction dynamics. Its control representation separates workspace, target objects, irrelevant objects, lighting, and background, enabling fine-grained scene edits across multiple robot embodiments. The training corpus includes 9.5M single-step samples and 2.6M sequential video clips, with Qwen3-VL-235B producing structured scene and trajectory annotations. For efficient decoding, FlashAR+ generates image tokens along anti-diagonals, reaching up to 82.9x faster 1024-by-1024 image generation than standard autoregressive decoding. Against GPT-Image-2.0, U0 wins human evaluations for embodied transfer and scene generation, while its synthetic visual variations raise π0.5 out-of-distribution task progress from 36.9% to 63.2% on real-world manipulation. On WorldArena, it ranks first among more than 100 submissions with an EWMScore of 73.64. The paper’s central claim is that a foundation world model can function not only as a predictive embodied simulator, but also as a scalable data engine for training more robust robot policies.
Original abstract
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.