ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
AuthorsHao Li, Ganlong Zhao, Yufei Liu, Haotian Hou, Guoquan Ye, Tongyan Fang, Chunxiao Liu, Siyuan Huang, Jianbo Liu, Xiaogang Wang, Hongsheng Li
Resources
ACE-Ego-0 makes robot learning more scalable by turning egocentric human videos into usable action supervision and training vision-language-action models on both human and robot data together.
Key results
Total heterogeneous embodied data used for ACE-Ego-0 pretraining
Sensor-logged robot demonstrations plus simulation rollouts
Egocentric human video converted into robot-format pseudo-actions
Average success rate of ACE-Ego-0
Average success rate of ACE-Ego-0
Average success rate across six real-world manipulation tasks
What the paper found
ACE-Ego-0, from ACE Robotics with CUHK MMLab and collaborators at CUHK Shenzhen, SJTU, and THU, is a Vision-Language-Action pretraining framework that unifies egocentric human video and multi-embodiment robot data by converting raw human footage into robot-format pseudo-actions, projecting both sources into a shared camera-space action representation, conditioning the policy with morphology tokens derived from URDFs or learned human surrogates, and aligning temporal supervision by physical duration instead of frame count. The method also introduces a reliability-aware objective that treats noisy human pseudo-actions as auxiliary supervision, emphasizing high-confidence position channels while down-weighting unstable rotation and gripper labels. Instantiated on a 6.0K+ hour mixed corpus, including 4.53K+ hours of robot and simulation data plus 1.48K hours of pseudo-action-labeled egocentric human video, and trained with Qwen3-VL-4B-Instruct plus a 600M flow-matching action expert, ACE-Ego-0 reaches 72.8% on RoboCasa GR1 TableTop, 91.12% on RoboTwin 2.0 Easy, 90.62% on RoboTwin 2.0 Hard, and 78.3% on a real ARX bimanual platform, outperforming π0.5 and GR00T-N1.7. Ablations show that removing human auxiliary reliability weighting causes the largest drop, from 72.8% to 69.2%, confirming that scale alone is not enough without label-quality control.
Original abstract
Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on camera-space actions, morphology conditioning, and time-aligned action chunking. To robustly leverage noisy pseudo-action supervision from egocentric human videos, we formulate a reliability-aware training objective with a human auxiliary loss that concentrates supervision on reliable signals. We instantiate ACE-EGO-0 on 4.53K hours of robot and simulation data, together with 1.48K hours of pseudo-action-labeled egocentric human data. Experiments show that incorporating large-scale human supervision under reliability-aware weighting consistently improves both unified joint pretraining and supervised fine-tuning. ACE-EGO-0 achieves state-of-the-art performance on RoboCasa GR1 TableTop and RoboTwin 2.0, while demonstrating strong transfer to real-world bimanual manipulation.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.