Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
AuthorsTara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
AffiliationsS. Shankar Sastry, Claire Tomlin*, and Jitendra Malik*
Resources
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.
Key results
Human hand-object trajectories used for evaluation across Dex3, Allegro, and Sharpa.
Percentage-point improvement in location-aware contact F1 over the strongest baseline.
Residual-RL task success rate in simulation using MMO references.
Frames per second on an NVIDIA RTX 4090, excluding inverse kinematics.
Success across 300 hardware trials involving 30 objects from 10 categories.
What the paper found
Morphometric Imitation converts reconstructed human hand-object interactions into zero-shot robot manipulation through three stages: Morphometric Optimization, or MMO, first scales the MANO hand model to match robot morphology and then recovers demonstrated contacts; residual reinforcement learning in ManiSkill adds dynamically feasible motion using object pose, contact state, PPO, and collision-aware termination; finally, demonstrations train a point-cloud visuomotor student with ManiFlow’s DiT-X action generator and flow-matching objectives. Across 10 GRAB trajectories and three robot hands—Dex3, Allegro, and Sharpa—MMO improves location-aware contact F1 over the strongest baseline by 27.8 percentage points on Dex3, while reducing contact patch error. The resulting residual policies reach success rates of 82.5% on Dex3, 69.9% on Allegro, and 91.8% on Sharpa. MMO runs at 180 frames per second on an NVIDIA RTX 4090, excluding inverse kinematics. For sim-to-real transfer, policies trained with randomized object scale, mass, friction, pose, and point-cloud observations achieve 89.3% success across 300 real-world trials involving 30 objects from 10 categories, without real-world policy training. The central finding is that morphology alignment and spatially precise contact preservation improve both dynamic feasibility and visual policy transfer, while object pose information guides task motion and contact information preserves the demonstrated grasp.
Original abstract
Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO improves contact F1 over the strongest of five baselines by at least 8 points for every hand, while also improving the success rate of downstream dynamic retargeting by as much as 35 points. Ablations on the residual RL show complementary benefits from using object pose and contact information. Finally, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects. Project page: $\href{https://morphometricimitation.github.io}{\text{this https URL}}$
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Agent as Policy for Robotic Manipulation
Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang
A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.