NTH

Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

AuthorsTara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik

AffiliationsS. Shankar Sastry, Claire Tomlin*, and Jitendra Malik*

October 4, 2026 2 min read
Watch on YouTube
The one-line take

A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.

Key results

10
GRAB trajectories

Human hand-object trajectories used for evaluation across Dex3, Allegro, and Sharpa.

27.8
MMO F1 gain on Dex3

Percentage-point improvement in location-aware contact F1 over the strongest baseline.

91.8%
Sharpa dynamic success

Residual-RL task success rate in simulation using MMO references.

180
MMO throughput

Frames per second on an NVIDIA RTX 4090, excluding inverse kinematics.

89.3%
Zero-shot real-world success

Success across 300 hardware trials involving 30 objects from 10 categories.

What the paper found

Morphometric Imitation converts reconstructed human hand-object interactions into zero-shot robot manipulation through three stages: Morphometric Optimization, or MMO, first scales the MANO hand model to match robot morphology and then recovers demonstrated contacts; residual reinforcement learning in ManiSkill adds dynamically feasible motion using object pose, contact state, PPO, and collision-aware termination; finally, demonstrations train a point-cloud visuomotor student with ManiFlow’s DiT-X action generator and flow-matching objectives. Across 10 GRAB trajectories and three robot hands—Dex3, Allegro, and Sharpa—MMO improves location-aware contact F1 over the strongest baseline by 27.8 percentage points on Dex3, while reducing contact patch error. The resulting residual policies reach success rates of 82.5% on Dex3, 69.9% on Allegro, and 91.8% on Sharpa. MMO runs at 180 frames per second on an NVIDIA RTX 4090, excluding inverse kinematics. For sim-to-real transfer, policies trained with randomized object scale, mass, friction, pose, and point-cloud observations achieve 89.3% success across 300 real-world trials involving 30 objects from 10 categories, without real-world policy training. The central finding is that morphology alignment and spatially precise contact preservation improve both dynamic feasibility and visual policy transfer, while object pose information guides task motion and contact information preserves the demonstrated grasp.

Original abstract

Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO improves contact F1 over the strongest of five baselines by at least 8 points for every hand, while also improving the success rate of downstream dynamic retargeting by as much as 35 points. Ablations on the residual RL show complementary benefits from using object pose and contact information. Finally, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects. Project page: $\href{https://morphometricimitation.github.io}{\text{this https URL}}$

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis
03Embodied Ai

Agent as Policy for Robotic Manipulation

Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang

A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.

Read analysis