NTH

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

AuthorsYuhang Huang, Xuan Lv, Junyan Xu, Zhiyuan Yu, Jiazhao Zhang, Ruizhen Hu, Wancheng Feng, Shilong Zou, Hewen Xiao, Ziqiao Zhou, Kaiyun Huang, Zhiyu Peng, Juzhan Xu, Hang Zhao, Chenyang Zhu, Renjiao Yi, Yifei Huang, Douhui Wu, Yan Zhang, Kexu Cheng, Chunhe Song, Yunzhi Xue, Xiuhong Zhang, Leitao Guo, Yunji Chen, Bin Wu, Haibin Yu, Kai Xu

June 26, 2026 2 min read
Watch on YouTube
The one-line take

PAIWorld upgrades world models for robots by making multiple camera views agree in 3D, improving how models simulate and plan manipulation tasks.

Key results

14B
parameters

Approximate size of the PAIWorld model built on Cosmos-Predict2.5

2.5M
training clips

Multi-view robotic manipulation video clips used for pretraining

72.31
WorldArena EWMScore

Best overall score on the WorldArena benchmark

0.8245
AgiBot-Challenge2026 EWMScore

Overall score on the AgiBot-Challenge2026 benchmark

14.20
AgiBot-World MEt3R

Best multi-view 3D reconstruction error on the AgiBot-World benchmark

What the paper found

PAIWorld, developed by the Institute of AI for Industries at the Chinese Academy of Sciences, extends a DiT-based world foundation model into a 3D-consistent multi-view simulator for robotic manipulation by combining two architectural pieces and one geometric training signal: Geometry-Aware Cross-View Attention, Geo-RoPE, and Latent 3D-REPA distilled from the frozen Depth Anything 3 model. Built on Cosmos-Predict2.5 with about 14B parameters and trained on 2.5M multi-view robot video clips from AgiBot-World, RoboMIND, Galaxea, RoboTwin, and RoboCOIN, it injects camera ray directions and extrinsics into attention while aligning token relations to 3D-aware features, reducing cross-view drift, depth conflict, and texture mismatch. On WorldArena, PAIWorld ranks 1st with an EWMScore of 72.31 and the best Motion Quality among all entries, and on AgiBot-Challenge2026 it ranks 2nd with an EWMScore of 0.8245 while achieving the best Scene Consistency at 0.9041. On the AgiBot-World text-conditioned benchmark, it reaches SSIM 0.7683, LPIPS 0.1844, FID 45.04, FVD 175.7778, and MEt3R 14.20, improving geometric reconstruction error by 10 percent over the next best method. Ablations show the full system cuts MEt3R from 16.84 to 14.20, a 2.64 gain that is larger than using cross-view attention or REPA alone, confirming that explicit inter-view communication and 3D prior supervision are both necessary.

Original abstract

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on multiple cameras (egocentric, eye-to-hand, and wrist-mounted) for policy learning, current multi-view world models simply concatenate view tokens without explicit geometric reasoning. This causes cross-view object drift, depth inconsistency, and texture misalignment. We trace these failures to two deficiencies: the absence of an explicit inter-view communication mechanism and the lack of a 3D geometric prior. We argue that resolving both simultaneously is necessary and sufficient. To address this, we present PAIWorld, a framework that augments diffusion-transformer world models via three core components: (1) Geometry-Aware Cross-View Attention blocks that establish an explicit pathway across views, (2) Geometric Rotary Position Embedding that encodes camera ray directions and extrinsic poses into the attention mechanism, and (3) Latent 3D-REPA, which distills 3D-aware features from frozen 3D foundation models to ensure 3D consistency. Built upon a DiT-based world foundation model, PAIWorld achieves state-of-the-art multi-view 3D consistency on robotic manipulation benchmarks, ranking 1st on the WorldArena leaderboard and 2nd on the AgiBot-Challenge2026 leaderboard, while enabling downstream applications such as model-based planning, world action models, and multi-view policy post-training.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis