NTH

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

AuthorsHao Li, Ganlong Zhao, Yufei Liu, Haotian Hou, Guoquan Ye, Tongyan Fang, Chunxiao Liu, Siyuan Huang, Jianbo Liu, Xiaogang Wang, Hongsheng Li

June 18, 2026 2 min read
Watch on YouTube
The one-line take

ACE-Ego-0 makes robot learning more scalable by turning egocentric human videos into usable action supervision and training vision-language-action models on both human and robot data together.

Key results

6.0K+
mixed pretraining pool

Total heterogeneous embodied data used for ACE-Ego-0 pretraining

4.53K+
robot and simulation data

Sensor-logged robot demonstrations plus simulation rollouts

1.48K
pseudo-action human video

Egocentric human video converted into robot-format pseudo-actions

72.8%
RoboCasa GR1 TableTop

Average success rate of ACE-Ego-0

91.12%
RoboTwin 2.0 Easy

Average success rate of ACE-Ego-0

78.3%
real ARX bimanual platform

Average success rate across six real-world manipulation tasks

What the paper found

ACE-Ego-0, from ACE Robotics with CUHK MMLab and collaborators at CUHK Shenzhen, SJTU, and THU, is a Vision-Language-Action pretraining framework that unifies egocentric human video and multi-embodiment robot data by converting raw human footage into robot-format pseudo-actions, projecting both sources into a shared camera-space action representation, conditioning the policy with morphology tokens derived from URDFs or learned human surrogates, and aligning temporal supervision by physical duration instead of frame count. The method also introduces a reliability-aware objective that treats noisy human pseudo-actions as auxiliary supervision, emphasizing high-confidence position channels while down-weighting unstable rotation and gripper labels. Instantiated on a 6.0K+ hour mixed corpus, including 4.53K+ hours of robot and simulation data plus 1.48K hours of pseudo-action-labeled egocentric human video, and trained with Qwen3-VL-4B-Instruct plus a 600M flow-matching action expert, ACE-Ego-0 reaches 72.8% on RoboCasa GR1 TableTop, 91.12% on RoboTwin 2.0 Easy, 90.62% on RoboTwin 2.0 Hard, and 78.3% on a real ARX bimanual platform, outperforming π0.5 and GR00T-N1.7. Ablations show that removing human auxiliary reliability weighting causes the largest drop, from 72.8% to 69.2%, confirming that scale alone is not enough without label-quality control.

Original abstract

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on camera-space actions, morphology conditioning, and time-aligned action chunking. To robustly leverage noisy pseudo-action supervision from egocentric human videos, we formulate a reliability-aware training objective with a human auxiliary loss that concentrates supervision on reliable signals. We instantiate ACE-EGO-0 on 4.53K hours of robot and simulation data, together with 1.48K hours of pseudo-action-labeled egocentric human data. Experiments show that incorporating large-scale human supervision under reliability-aware weighting consistently improves both unified joint pretraining and supervised fine-tuning. ACE-EGO-0 achieves state-of-the-art performance on RoboCasa GR1 TableTop and RoboTwin 2.0, while demonstrating strong transfer to real-world bimanual manipulation.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis