GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
AuthorsTianyi Xie, Haotian Zhang, Jinhyung Park, Zi Wang, Bowen Wen, Jiefeng Li, Xueting Li, Qingwei Ben, Haoyang Weng, Yufei Ye, David Minor, Tingwu Wang, Chenfanfu Jiang, Sanja Fidler, Jan Kautz, Linxi Fan, Yuke Zhu, Zhengyi Luo, Umar Iqbal, Ye Yuan
Resources
GRAIL uses 3D scenes plus video-model priors to generate large-scale humanoid interaction data in simulation, then turns that synthetic data into real robot skills like picking up objects and climbing stairs.
Key results
GRAIL dataset size spanning pick-up, whole-body manipulation, sitting, and terrain traversal
3D object assets used to generate the loco-manipulation dataset
Physical executability on the shared 20-object evaluation set
Real-world Unitree G1 success rate for diverse object pick-up
Real-world Unitree G1 success rate for stair-climbing
What the paper found
GRAIL, from NVIDIA and UCLA, is a fully digital pipeline for generating humanoid loco-manipulation from 3D assets and video priors, designed to avoid physical scene rebuilds, teleoperation, and motion capture until deployment. It first constructs a metric 3D scene with known camera, scale, object geometry, and a robot-proportioned character, then uses video foundation models and OpenAI’s ChatGPT-assisted prompting to synthesize interaction videos, reconstructs metric 4D human-object trajectories with GENMO, WiLoR, FoundationPose, MoGe-2, SAM2, and interaction-aware losses for projection, depth, and contact, and finally retargets motions to a Unitree G1 through task-general trackers built on the SONIC whole-body controller. The resulting asset-conditioned dataset contains over 20,000 sequences spanning pick-up, whole-body manipulation, sitting, and terrain traversal, generated from 1,000 object assets and 1,000 terrain configurations. On a shared 20-object benchmark, GRAIL reaches 0.008 contact distance, 0.90% penetration, a 3.58 interaction score, and 88.9% physical executability success rate, outperforming DAViD, CHOIS, and HOIDiff. For downstream control, the full tracking system achieves 81.4% success with 0.135 object-position error and 41.8 local MPJPE, and its sim-to-real egocentric policies transfer to hardware with 84% pick-up success on diverse objects and 90% stair-climbing success on a real Unitree G1.
Original abstract
Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scale because each collection depends on physical setups, instrumented actors, and robot operation. We present GRAIL, a digital generation pipeline that remains fully virtual until deployment: it composes 3D assets, simulator-ready scenes, and priors from video foundation models (VFMs) to synthesize interactions without rebuilding physical environments or teleoperating the robot. Rather than reconstructing unconstrained in-the-wild videos, GRAIL starts from fully specified 3D configurations in which object geometry, camera parameters, metric scale, environment depth, and a robot-proportioned character are known before video generation and reused during reconstruction. This privileged setup better conditions 4D recovery, allowing model-based object tracking, human motion estimation, and interaction-aware optimization to reconstruct metric 4D human-object interaction (HOI) trajectories with reduced depth ambiguity and morphology mismatch. We retarget the recovered motions to a humanoid robot and train complementary task-general trackers: an object-aware latent adaptor for manipulation and a scene-aware tracker for terrain traversal. GRAIL produces over 20,000 sequences spanning pick-up, object manipulation, sitting, and terrain traversal. Using only GRAIL-generated data, we train egocentric visual policies through a sim-to-real pipeline and deploy them on a Unitree G1 humanoid, achieving 84\% real-world success on diverse object pick-up and 90\% success on stair-climbing.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.