One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry
AuthorsSatvik Sharma, Samrat Sahoo, Huang Huang, Fei-Fei Li Jiajun Wu, Dorsa Sadigh, Jeannette Bohg
Resources
DemoMimic teaches robot hands to generalize dexterous manipulation by concentrating on the local geometry where contact occurs.
Key results
Average success across 16 objects, 4 tasks, and 2 robot-hand embodiments.
Unique real-world objects used to test generalization.
Manipulation tasks evaluated in the real world.
Success rate across tasks using the Sharpa embodiment.
Success rate across tasks using the Tesollo embodiment.
Open-the-box performance when local contact geometry differs.
What the paper found
DemoMimic learns dexterous manipulation from a single human demonstration and transfers it across unseen objects by exploiting local contact geometry rather than matching complete object shape. The framework combines residual reinforcement learning with 2 contact-centric rewards: Alignment Reward, which aligns hand and object surface normals, and Sustained Contact Reward, which favors uninterrupted contact streaks. A high-level Diffusion Policy predicts a wrist trajectory from one egocentric RGB image, while a low-level Diffusion Policy uses wrist-mounted depth, proprioception, and the trajectory for closed-loop control without explicit object-pose estimation. Demonstrations come from the ARCTIC dataset or manually captured motion using Meta Quest tracking; simulation imagery is diversified with NVIDIA’s Cosmos-Transfer-2.5, and deployment depth is produced by FoundationStereo. On 16 objects spanning 4 tasks and 2 robot-hand embodiments, the system reaches 71% real-world success, with 76% on the Sharpa hand and 65% on the Tesollo hand. Generalization fails mainly when local contact geometry changes: the robot-hand box reaches only 39% on the open-the-box task. Ablations show that removing either contact reward increases hardware failures, despite near-ceiling simulation performance, demonstrating that precise alignment and persistent forceful contact are central to sim-to-real transfer.
Original abstract
Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-scale dexterous robot hand data remains difficult. Learning from human demonstrations has emerged as a scalable alternative to robot teleoperation, providing strong priors on object interaction and contact strategies. Recent sim-to-real RL methods incorporate such priors, but often (i) omit rewards that explicitly incentivize precise contact, yielding weak real-world performance, and/or (ii) generalize poorly to unseen object instances. We propose DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points. Its contact-centric rewards encourage precise contact and improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved. Real-world ablations show that DemoMimic achieves 71% success across 16 objects, four tasks, and two robot-hand embodiments, with the smallest sim-to-real drop compared to baselines.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.