NTH

Human Universal Grasping

AuthorsKevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto

June 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches robots to grasp like humans by learning from a million egocentric human grasps and using that data to generate and retarget natural grasps in the real world.

Key results

1M
1M-HUG S frames

egocentric training frames in the human grasp dataset

6707
6,707 recordings

collection sessions across indoor environments

41
41 buildings

distinct buildings covered by the dataset

90
HUG-B ENCH objects

unseen evaluation objects with metric-scale meshes

73.0%
test success rate

full RGB+PC HUG test performance in simulation

66.7%
real-world tabletop success

HUG success on the 30-object tabletop test split

What the paper found

Human Universal Grasping, or HUG, from New York University, Tsinghua University, the University of Michigan, and Meta Reality Labs’ Project Aria team, reframes dexterous grasp learning around human data instead of robot teleoperation or simulation. The system is trained on 1M-HUG S, a 1M-frame egocentric dataset collected with Aria Gen 2 smart glasses across 6,707 recordings and 41 buildings, then learns a flow-matching transformer that fuses RGB-D with a queried 3D object point to predict a 99-dimensional MANO grasp state: wrist translation, 6D wrist rotation, and 15 joint rotations. HUG-B ENCH standardizes evaluation with 90 unseen objects spanning five geometry classes and three size bins, reconstructed into metric-scale meshes for both simulation and real-robot trials. In MuJoCo, the full RGB+PC model reaches 73.0% success on the test split with 14.6 mm fingertip contact error, while an oracle replay of recorded human grasps reaches 94.0% and 7.4 mm. The key ablation shows that removing the 3D supervision drops test success to 32.7%, and removing RGB drops the model to 70.7%, confirming that geometry and semantics are complementary. In real-world tabletop trials on 30 unseen objects, HUG achieves 66.7% success, outperforming Dex1B at 43.7% and CAP at 32.7%, and it still reaches 62.0% in an uncontrolled in-the-wild household setting after zero-shot retargeting to the Ability Hand and WUJI Hand without per-hand training.

Original abstract

Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up thousands of objects every day. We present HUG, a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. Using smart glasses, we first collect 1M-HUGs, an egocentric dataset of human grasps spanning 1M frames (27.8 hrs) and 6,707 object instances across 41 buildings. Next, to model the distribution of natural human grasps, our novel flow-matching model fuses RGB and depth observations to output a grasp parameterized by wrist translation, wrist rotation, and MANO hand pose. Predicted grasps can be retargeted to various robot hands, enabling zero-shot grasping in everyday scenes. To standardize evaluation, we build a new simulated benchmark, HUG-Bench, of 90 unseen objects from five geometric categories and various sizes, with metric-scale 3D meshes. We evaluate HUG in the real world on the 30-object test set of HUG-Bench across multiple stereo cameras, robot embodiments, and household environments. HUG outperforms the state-of-the-art grasping baselines by +23% and +34% on our challenging object set. Code, data, benchmark, checkpoints, and an interactive demo are released on our website: https://grasping.io/

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis