NTH

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

AuthorsYukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu

August 10, 2026 2 min read
Watch on YouTube
The one-line take

ACE-Data-0 turns everyday homes into multisensory recording studios for teaching embodied AI how humans see, move, touch, and complete long-horizon tasks.

Key results

150
ACE-Data-0 duration

Hours of synchronized multimodal home interaction capture.

17M
ACE-Data-0 video frames

Total video frames released in the dataset.

75000
ACE-Data-0 interaction episodes

Goal-directed interaction episodes across the captured activities.

0.7095
TouchAnything temporal accuracy

Best reported temporal accuracy for tactile inference from egocentric video.

9.1
WiLoR PA-MPJPE

Millimeter hand-pose error for exocentric hand reconstruction.

What the paper found

ACE-Data-0 addresses embodied AI’s data bottleneck with the Ambient Capture Engine, a two-scale system that turns real homes into synchronized recording studios. A table-scale setup resolves dexterous hand-object manipulation, while a room-scale setup captures locomotion, whole-body motion, and human-scene interaction. The system aligns egocentric and exocentric video, 41-joint body motion, articulated hand pose, object 6-DoF trajectories, audio, and tactile pressure in a shared spatial-temporal frame, using optical-clock synchronization with millisecond-level residual error and NVIDIA Jetson Orin hardware for multi-camera ingest. The resulting ACE-Data-0 contains 150 hours, 17M video frames, 75,000 interaction episodes, 200 task categories, 50 participants, and 2 environments. Goal-level instructions preserve natural planning, hesitation, improvisation, and long-horizon household behavior rather than scripted action sequences. Its hierarchical benchmark evaluates tactile inference, human-motion recovery, and hand-motion estimation. TouchAnything reaches 0.7095 temporal accuracy, but only 0.1646 contact IoU, exposing difficulty in recovering precise pressure under occlusion. For exocentric hand reconstruction, WiLoR achieves 9.1 mm PA-MPJPE, while HaPTIC reports 63 mm trajectory error; egocentric world-space methods remain near 100 mm. Gemini-3.1-pro-preview generates aligned language descriptions, making the dataset suitable for imitation learning, world models, and vision-language-action systems.

Original abstract

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis