Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
AuthorsYe Wang, Pei Lin, Xiong-Hui Chen, Haoqi Yuan, Zhixuan Liang, Yiyang Huang, Anzhe Chen, Zixing Lei, Jie Zhang, Tao Zhang, Haoyang Li, Tong Zhang, Chenxi Xiao, Ziyuan Jiao, Qin Jin
Resources
Ego2Robot turns everyday human manipulation videos into large-scale synthetic robot experience to help robots generalize beyond their training environments.
Key results
Hours of Ego2Robot training data generated across 15 robot morphologies.
Hours processed from ANT, EgoDex, ViTRA, and EgoVerse before robot rendering.
Success rate for 1:1 Ego2Robot and robot-data pretraining.
Percentage-point gain over the 50.9% robot-only baseline.
Percentage-point gain achieved by the 3:1 Ego2Robot-to-robot mix, reaching 51.7%.
Percentage-point improvement over robot-only training on the ARX ACone platform.
What the paper found
Ego2Robot, from researchers at Renmin University of China, Alibaba’s Qwen Team, ShanghaiTech, and BIGAI, presents a pipeline for turning egocentric human manipulation video into robot-training data at scale. It estimates hand motion with WiLoR and DynHaMR, retargets trajectories to robot grippers, searches feasible robot bases with MuJoCo inverse kinematics, removes human arms using SAM 3 and ProPainter, composites rendered robots with depth, and applies trajectory, statistical, and Qwen3.5 vision-language consistency filtering. From 1,940 hours across ANT, EgoDex, ViTRA, and EgoVerse, the system produces 18,561 hours across 15 robot morphologies. Using a Qwen3.5-4B vision-language-action model with a Diffusion Transformer action head, the authors extend RoboTwin2.0 to isolate visual, scene, embodiment, and task-semantic shifts. Mixing Ego2Robot and real robot data at 1:1 reaches 53.5% on RoboTwin Randomized, 2.6 percentage points above robot-only pretraining, while the 3:1 mix reaches 51.7% on EBench, a 12.1-point gain. Improvements are strongest for lighting, camera offsets, cross-embodiment transfer, unseen objects, and paraphrased instructions. In an ARX ACone deployment, adding converted egocentric play data improves five long-horizon tasks, with gains of 14 points on Put Blocks and 13 points on Insert Screw, suggesting that synthesized human demonstrations complement rather than replace real robot data.
Original abstract
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present \textbf{Ego2Robot}, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.