HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
AuthorsJuncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai, Kewei Zhang, Ye Huang, Bo Liang, Shukai Gong, Jiankai Tu, Xiaotian Tang, Jiaxin Li, Kaiqi Chen, Duomin Wang, Yuqi Wang, Bingyi Kang, Eric Huang, Zhiyang Dou, Zhen Dong, Enze Xie, Wojciech Matusik, Tat-Seng Chua, Daquan Zhou
Resources
This study suggests a surprising shortcut for embodied AI: pretrain on cheap, abundant human egocentric video, then fine-tune on a small amount of robot data for the best real-world performance.
Key results
Matched pretraining size for egocentric human video and real-robot data
AgiBot World manipulation post-training set used for adaptation
Best validation action loss with egocentric pretraining on in-distribution tasks
Best validation action loss with egocentric pretraining on out-of-distribution tasks
Best validation action loss with matched real-robot pretraining on out-of-distribution tasks
Egocentric-pretrained policy success rate in distribution on AgiBot bimanual rollouts
What the paper found
HumanScale, from PKU, NUS, MIT, UCSB, and NVIDIA, tests a simple but consequential question for embodied pretraining: can egocentric human video beat teleoperated robot data when both are matched for scale? Using the same autoregressive World-Action Model built on Wan 2.2 with a Mixture-of-Transformers architecture, the authors pretrain on 5,000 hours of curated HumanNet egocentric video or 5,000 hours of real-robot trajectories, then post-train on the same 1,500-trajectory AgiBot World manipulation set and evaluate on held-out seen and unseen tasks. The result is that human video not only scales cleanly but generalizes better: validation action loss falls from 0.0080 to 0.0067 on seen tasks and from 0.0234 to 0.0204 on unseen tasks as egocentric pretraining grows from 100 to 5,000 hours, following a strong log-linear law, while matched robot pretraining stalls on unseen tasks near 0.0254. Relative to no pretraining, egocentric data yields a 35% lower seen loss and 24% lower unseen loss; compared with robot pretraining at 5,000 hours, it is about 20% better on unseen tasks. The strongest evidence comes in real-robot rollouts on an AgiBot bimanual platform: the egocentric-pretrained policy reaches 92.5% success in distribution and 90.0% out of distribution, whereas the no-pretraining baseline collapses from 40.0% to 0.0%. The paper’s core claim is that egocentric video supplies the coverage and diversity needed for pretraining, while a small amount of robot data is sufficient later to close the embodiment gap.
Original abstract
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and validation protocols. Surprisingly, we find that egocentric data, when processed through a carefully designed filtering and labeling pipeline, is not merely a viable substitute for model pretraining but can lead to superior performance. With the same amount of pretraining data, models pretrained on egocentric data achieve a 24% lower validation loss on real-robot action prediction, as well as 52.5% and 90% higher success rates on in-distribution and out-of-distribution real-robot task execution, respectively. This finding verifies a scalable paradigm for embodied foundation models: pretrain on egocentric human video to learn diverse world representations, then adapt with a small amount of labeled real-robot data for action-space alignment. We hope this study encourages broader exploration of egocentric data and offers guidance for data quality assessment before costly robot data collection.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.