ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
AuthorsYukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu
Resources
ACE-Data-0 turns everyday homes into multisensory recording studios for teaching embodied AI how humans see, move, touch, and complete long-horizon tasks.
Key results
Hours of synchronized multimodal home interaction capture.
Total video frames released in the dataset.
Goal-directed interaction episodes across the captured activities.
Best reported temporal accuracy for tactile inference from egocentric video.
Millimeter hand-pose error for exocentric hand reconstruction.
What the paper found
ACE-Data-0 addresses embodied AI’s data bottleneck with the Ambient Capture Engine, a two-scale system that turns real homes into synchronized recording studios. A table-scale setup resolves dexterous hand-object manipulation, while a room-scale setup captures locomotion, whole-body motion, and human-scene interaction. The system aligns egocentric and exocentric video, 41-joint body motion, articulated hand pose, object 6-DoF trajectories, audio, and tactile pressure in a shared spatial-temporal frame, using optical-clock synchronization with millisecond-level residual error and NVIDIA Jetson Orin hardware for multi-camera ingest. The resulting ACE-Data-0 contains 150 hours, 17M video frames, 75,000 interaction episodes, 200 task categories, 50 participants, and 2 environments. Goal-level instructions preserve natural planning, hesitation, improvisation, and long-horizon household behavior rather than scripted action sequences. Its hierarchical benchmark evaluates tactile inference, human-motion recovery, and hand-motion estimation. TouchAnything reaches 0.7095 temporal accuracy, but only 0.1646 contact IoU, exposing difficulty in recovering precise pressure under occlusion. For exocentric hand reconstruction, WiLoR achieves 9.1 mm PA-MPJPE, while HaPTIC reports 63 mm trajectory error; egocentric world-space methods remain near 100 mm. Gemini-3.1-pro-preview generates aligned language descriptions, making the dataset suitable for imitation learning, world models, and vision-language-action systems.
Original abstract
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.