Human Universal Grasping
AuthorsKevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
Resources
This paper teaches robots to grasp like humans by learning from a million egocentric human grasps and using that data to generate and retarget natural grasps in the real world.
Key results
egocentric training frames in the human grasp dataset
collection sessions across indoor environments
distinct buildings covered by the dataset
unseen evaluation objects with metric-scale meshes
full RGB+PC HUG test performance in simulation
HUG success on the 30-object tabletop test split
What the paper found
Human Universal Grasping, or HUG, from New York University, Tsinghua University, the University of Michigan, and Meta Reality Labs’ Project Aria team, reframes dexterous grasp learning around human data instead of robot teleoperation or simulation. The system is trained on 1M-HUG S, a 1M-frame egocentric dataset collected with Aria Gen 2 smart glasses across 6,707 recordings and 41 buildings, then learns a flow-matching transformer that fuses RGB-D with a queried 3D object point to predict a 99-dimensional MANO grasp state: wrist translation, 6D wrist rotation, and 15 joint rotations. HUG-B ENCH standardizes evaluation with 90 unseen objects spanning five geometry classes and three size bins, reconstructed into metric-scale meshes for both simulation and real-robot trials. In MuJoCo, the full RGB+PC model reaches 73.0% success on the test split with 14.6 mm fingertip contact error, while an oracle replay of recorded human grasps reaches 94.0% and 7.4 mm. The key ablation shows that removing the 3D supervision drops test success to 32.7%, and removing RGB drops the model to 70.7%, confirming that geometry and semantics are complementary. In real-world tabletop trials on 30 unseen objects, HUG achieves 66.7% success, outperforming Dex1B at 43.7% and CAP at 32.7%, and it still reaches 62.0% in an uncontrolled in-the-wild household setting after zero-shot retargeting to the Ability Hand and WUJI Hand without per-hand training.
Original abstract
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up thousands of objects every day. We present HUG, a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. Using smart glasses, we first collect 1M-HUGs, an egocentric dataset of human grasps spanning 1M frames (27.8 hrs) and 6,707 object instances across 41 buildings. Next, to model the distribution of natural human grasps, our novel flow-matching model fuses RGB and depth observations to output a grasp parameterized by wrist translation, wrist rotation, and MANO hand pose. Predicted grasps can be retargeted to various robot hands, enabling zero-shot grasping in everyday scenes. To standardize evaluation, we build a new simulated benchmark, HUG-Bench, of 90 unseen objects from five geometric categories and various sizes, with metric-scale 3D meshes. We evaluate HUG in the real world on the 30-object test set of HUG-Bench across multiple stereo cameras, robot embodiments, and household environments. HUG outperforms the state-of-the-art grasping baselines by +23% and +34% on our challenging object set. Code, data, benchmark, checkpoints, and an interactive demo are released on our website: https://grasping.io/
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.