ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
AuthorsJayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa
Resources
ADEPT teaches dexterous robots reusable manipulation skills through RL pretraining, enabling them to transfer complex visual-tactile behaviors to new tasks and real-world hardware.
Key results
ADEPT uses 11B total environment steps, including pretraining and downstream post-training.
Each new downstream task requires 3B post-training steps after the reusable pretraining cost.
Flexiv-Sharpa achieves 8/10 real-world FMB square/round insertions with five fingertip tactile sensors.
The vision-only counterpart achieves 3/10 insertions on the comparable FMB evaluation.
Complete manipulation sequences execute in 5–10 seconds per trial.
The FMB parallel-jaw pipeline requires 20–70 seconds per trial.
What the paper found
NVIDIA’s ADEPT, or Accelerating Dexterity via Pre-Training and Post-Training, applies PPO reinforcement learning to high-DoF robot hands through a reusable dexterity prior. The policy first learns reaching, grasping, lifting, in-hand reorientation, transport, and reposing across 16 randomized primitive shapes using ADR and Population-Based Training; downstream adaptation then combines 40k iterations of behavior-cloning actor distillation, 20 PPO iterations of critic warm-up, and conservative PPO updates that reduce the actor learning rate from 1e-3 to 1e-5 and tighten the clip from 0.20 to 0.05. A full joint-space Geometric Fabric enforces collision and joint-limit safety while exposing the complete 23-DoF Kuka-Allegro or 29-DoF Flexiv-Sharpa action space. Teacher policies are distilled with DAgger into students using two RGB cameras, an 8-keypoint pose auxiliary loss, and, on Flexiv-Sharpa, five TacMap fingertip tactile sensors. ADEPT reaches the FMB insertion endpoint using 11B total environment steps, with only 3B required for each new downstream task after pretraining. In zero-shot real-robot trials, the visuo-tactile Flexiv-Sharpa system achieves 8/10 insertions versus 3/10 for the vision-only counterpart, while complete manipulations take 5–10 s per trial compared with 20–70 s for the FMB parallel-jaw pipeline, a 2×–14× speed advantage.
Original abstract
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies with this pretrained behavior as a prior. ADEPT enables learning new behaviors that are otherwise difficult to discover from scratch on multi-fingered robots and avoids learning the same set of skills over again for every new downstream task. The pretrained policy zero-shots the reposing phase of downstream tasks, but naïve RL fine-tuning rapidly degrades this capability during transfer. We address this with a stable post-training recipe combining behavior-cloning distillation, critic warm-up, and conservative on-policy updates. To safely exploit the full kinematic dexterity, we introduce a joint-space Geometric Fabric that mediates between the RL policy and the robot. We distill post-trained teachers into perceptive students that zero-shot sim-to-real transfer on two embodiments: a 23 DoF Kuka-Allegro with two RGB cameras, and a 29 DoF Flexiv-Sharpa with two RGB cameras and five vision-based tactile sensors, and can solve long-horizon tasks from challenging initial states with dexterity at human-level speed.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.