NTH

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment

AuthorsXPolicyLab Community, Tianxing Chen, Yue Chen, Tian Nian, Zijian Cai, Guangyu Chen, Wenwei Lin, Qiwei Liang, Peicheng Xiang, Kailun Su, Zixuan Li, Junyuan Tang, Yan Qin, Qiangyu Chen, Shaolong Zhu, Xiang Li, Jiahao Zhang, Weijie Wan, Baijun Chen, Honghao Su, Kehe Ye, Shujia Liu, Kaixuan Wang, Haotian Liang, Yunze Liu, Mingleyang Li, Yuran Wang, Boyu Chen, Hongzhe Bi, Shuhe Huang, Hengkai Tan, Jisong Cai, Yao Mu, Jun Guo, Xiaofeng Wang, Zheng Zhu, Weijie Ke, Hengtao Li, Yuhang Tang, Xiaofan Li, Ganlin Yang, Zhangzheng Tu, Shuai Yang, Wenxuan Song, Pengxiang Ding, Kaidong Zhang, Yu Sun, Junliang Guo, Tong Zhang, Yixing Chen, Rongxu Cui, Zongzheng Zhang, Haoxiang Ma, Junhao Cai, Haoyu Zhang, Senqiao Yang, Jinhui Ye, Pengguang Chen, Shu Liu, Xiu Su, Wenhan Fang, Wenhao Li, Yichao Cao, Chengyao Wang, Qiang Chen, Ping Luo, Wenbo Ding

August 19, 2026 2 min read
Watch on YouTube
The one-line take

XPolicyLab aims to make evaluating and deploying robot policies as plug-and-play as using a common software standard instead of building a new integration for every policy and environment.

Key results

42
Integrated robot policies

Policies supported across VLA, world-action, diffusion, memory, and imitation-learning families.

5
From-scratch integration time

Hours required to connect π0.5 to RoboDojo from scratch.

2
Manual XPolicyLab integration time

Hours required when following the standard manually.

30
Agent-assisted integration time

Minutes required using machine-readable agent skills.

77.8%
FastWAM clean RoboTwin success rate

Highest clean-setting success rate in the reported RoboTwin leaderboard.

1.9%
FastWAM randomized RoboTwin success rate

FastWAM success rate under RoboTwin randomization.

What the paper found

XPolicyLab introduces a unified standard for evaluating and deploying heterogeneous robot policies, replacing the pairwise O(NM) integration problem with reusable policy adapters and environment clients whose scaling is O(N+M). Its contract standardizes multimodal observations, joint- and end-effector actions, stateful inference, action chunks, batching, and episode resets without constraining model architectures or training objectives. A dependency-isolated policy server communicates with simulator or robot clients through WebSocket RPC and MessagePack, allowing native stacks for models such as OpenVLA, π0.5, and GR00T-N1.7 to run locally or on remote GPUs while environments retain their own simulator and control software. The ecosystem integrates 42 policies across VLA, world-action, diffusion, memory-augmented, and imitation-learning families, and the same adapters operate across RoboTwin, RoboDojo simulation, and RoboDojo-RealEval. In a within-subject study with 6 engineers, connecting π0.5 to RoboDojo required over 5 hours from scratch, 2 hours with the manual standard, and 30 minutes when machine-readable agent skills guided Cursor with Opus 5; the skills also support coding agents such as Claude Code and Codex. On RoboTwin, FastWAM achieved the highest clean success rate at 77.8%, but only 1.9% under randomization, highlighting that the infrastructure improves reproducibility and portability rather than policy capability itself.

Original abstract

Robot policy evaluation and deployment remain fragmented by model-specific software dependencies, data representations, and runtime interfaces, so that connecting N policies to M evaluation environments requires O(NM) separate integrations. We present XPolicyLab, a unified standard and open ecosystem that reduces this cost to O(N+M). XPolicyLab specifies common observation, action, and trajectory schemas together with a minimal adapter interface for observation updates, action prediction, batched execution, and episode reset, while a dependency-isolated client/server architecture separates policy inference from environment execution, so that each side retains its native software stack and may run locally or remotely. The ecosystem integrates 42 robot policies and standardizes their installation, debugging, serving, and evaluation workflows. Across these adapters, model-specific code varies by an order of magnitude while the environment-facing loop stays within a few lines of a fixed reference, confirming that the contract confines heterogeneity to the policy side. In a controlled study, conforming to the standard reduces the integration effort of a representative policy from over five hours to two hours, and packaged agent skills reduce it further to thirty minutes. The same adapters serve RoboTwin, RoboDojo simulation, and standardized real-robot evaluation through one interface. XPolicyLab is released as shared infrastructure for reproducible policy comparison and standardized deployment across simulation and physical platforms. Project website: https://xpolicylab.github.io/.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis