CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
AuthorsJulien Merand, Boris Meden, Liming Chen, Mathieu Grossard
Resources
CoToGrasp learns to generate functionally meaningful dexterous grasps across unseen objects by conditioning on contact topology rather than object-specific annotations.
Key results
Semantic grasp conditions modeled by the planner.
Object-free examples generated from hand configurations and topology templates.
CoToGrasp score on DexGraspNet, compared with 14.28% for Dexonomy.
CoToGrasp entropy on taxonomy-aware evaluation, compared with 0.77 for Dexonomy.
Average seconds required to generate one grasp.
Physical success rate retained from convex to non-convex objects.
What the paper found
CoToGrasp is an object-agnostic generative planner for dexterous hands that conditions grasp synthesis on 21 semantic contact topologies, including precision, power, and object-specific patterns, rather than optimizing stability alone. It generates 210,000 training examples from 10,000 valid hand configurations without object meshes, then transfers local geometry into a gripper-centered canonical workspace using a DGCNN encoder and k-nearest-neighbor aggregation. A Transformer, Set Transformer, and conditional variational autoencoder predict topology-specific contact masks, while label-consistency filtering, force-closure checks, and energy-based joint optimization enforce kinematic feasibility, collision avoidance, and a 5-millimeter safety margin. On the DexGraspNet benchmark, CoToGrasp reaches 17.18% topology compliance versus 14.28% for Dexonomy, with semantic entropy of 0.84 versus 0.77, and generates a grasp in 0.11 seconds. It also retains 58.72% of its physical success rate when moving from convex to non-convex objects, compared with 37.42% for Dexonomy. The method was validated on an Allegro Hand mounted on a UR10 using YCB objects, and its implementation was trained with NVIDIA A100 GPUs, demonstrating zero-shot transfer to unseen geometries while preserving functional contact structure.
Original abstract
Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks. However, conditioning grasp synthesis on specific human grasp taxonomies typically requires prohibitively expensive, object-annotated datasets. To address these limitations, we propose CoToGrasp, a novel generative framework that synthesizes diverse, stable grasps strictly conditioned on specific contact topologies. To bypass the data collection bottleneck, CoToGrasp is trained entirely in an object-agnostic manner. We introduce a feature-based canonical workspace that projects local object features into a unified gripper-centric domain, effectively decoupling the semantic functional intent from the arbitrary object geometry. By learning the intrinsic contact manifold of the gripper within this workspace, our model achieves zero-shot generalization to unseen objects at inference. Extensive evaluations on the large-scale DexGraspNet dataset demonstrate that CoToGrasp achieves state-of-the-art performance, outperforming existing taxonomy-guided planners. Finally, we demonstrate the physical viability and kinematic feasibility of our synthesized contact topologies on a physical robot platform. Code is available on our project website https://cea-list.github.io/cotograspweb/ .
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.