NTH

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

AuthorsXinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li

July 21, 2026 2 min read
Watch on YouTube
The one-line take

Xiaomi-Robotics-U0 uses a large multimodal world foundation model to generate consistent robot-centered scenes and videos while improving real-world manipulation performance.

Key results

38B
Model size

Multimodal autoregressive parameters in Xiaomi-Robotics-U0

9.5M
Single-step dataset

Samples used for general and embodied one-step generation

2.6M
Sequential video dataset

Video clips used for sequential embodied generation

82.9
Image-generation acceleration

Maximum speedup from FlashAR-style decoding over standard autoregressive generation

63.2%
Out-of-distribution policy progress

π0.5 progress after augmentation, up from 36.9%

73.64
WorldArena EWMScore

Overall score, ranking first among more than 100 submissions

What the paper found

Xiaomi Robotics introduces Xiaomi-Robotics-U0, a 38B-parameter multimodal autoregressive world model that unifies text-to-image generation, image editing, multi-view embodied scene synthesis, structured embodied transfer, and robot video prediction. Initialized from EMU3.5 and trained with a shared next-token objective, it preserves general visual knowledge while adding robot-specific geometry, camera consistency, embodiment constraints, and interaction dynamics. Its control representation separates workspace, target objects, irrelevant objects, lighting, and background, enabling fine-grained scene edits across multiple robot embodiments. The training corpus includes 9.5M single-step samples and 2.6M sequential video clips, with Qwen3-VL-235B producing structured scene and trajectory annotations. For efficient decoding, FlashAR+ generates image tokens along anti-diagonals, reaching up to 82.9x faster 1024-by-1024 image generation than standard autoregressive decoding. Against GPT-Image-2.0, U0 wins human evaluations for embodied transfer and scene generation, while its synthetic visual variations raise π0.5 out-of-distribution task progress from 36.9% to 63.2% on real-world manipulation. On WorldArena, it ranks first among more than 100 submissions with an EWMScore of 73.64. The paper’s central claim is that a foundation world model can function not only as a predictive embodied simulator, but also as a scalable data engine for training more robust robot policies.

Original abstract

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis