NTH

ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning

AuthorsJayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa

August 28, 2026 2 min read
Watch on YouTube
The one-line take

ADEPT teaches dexterous robots reusable manipulation skills through RL pretraining, enabling them to transfer complex visual-tactile behaviors to new tasks and real-world hardware.

Key results

11B
Total training steps

ADEPT uses 11B total environment steps, including pretraining and downstream post-training.

3B
Marginal downstream training

Each new downstream task requires 3B post-training steps after the reusable pretraining cost.

8/10
Visuo-tactile real success

Flexiv-Sharpa achieves 8/10 real-world FMB square/round insertions with five fingertip tactile sensors.

3/10
Vision-only real success

The vision-only counterpart achieves 3/10 insertions on the comparable FMB evaluation.

5–10
ADEPT execution time

Complete manipulation sequences execute in 5–10 seconds per trial.

20–70
Parallel-jaw execution time

The FMB parallel-jaw pipeline requires 20–70 seconds per trial.

What the paper found

NVIDIA’s ADEPT, or Accelerating Dexterity via Pre-Training and Post-Training, applies PPO reinforcement learning to high-DoF robot hands through a reusable dexterity prior. The policy first learns reaching, grasping, lifting, in-hand reorientation, transport, and reposing across 16 randomized primitive shapes using ADR and Population-Based Training; downstream adaptation then combines 40k iterations of behavior-cloning actor distillation, 20 PPO iterations of critic warm-up, and conservative PPO updates that reduce the actor learning rate from 1e-3 to 1e-5 and tighten the clip from 0.20 to 0.05. A full joint-space Geometric Fabric enforces collision and joint-limit safety while exposing the complete 23-DoF Kuka-Allegro or 29-DoF Flexiv-Sharpa action space. Teacher policies are distilled with DAgger into students using two RGB cameras, an 8-keypoint pose auxiliary loss, and, on Flexiv-Sharpa, five TacMap fingertip tactile sensors. ADEPT reaches the FMB insertion endpoint using 11B total environment steps, with only 3B required for each new downstream task after pretraining. In zero-shot real-robot trials, the visuo-tactile Flexiv-Sharpa system achieves 8/10 insertions versus 3/10 for the vision-only counterpart, while complete manipulations take 5–10 s per trial compared with 20–70 s for the FMB parallel-jaw pipeline, a 2×–14× speed advantage.

Original abstract

We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies with this pretrained behavior as a prior. ADEPT enables learning new behaviors that are otherwise difficult to discover from scratch on multi-fingered robots and avoids learning the same set of skills over again for every new downstream task. The pretrained policy zero-shots the reposing phase of downstream tasks, but naïve RL fine-tuning rapidly degrades this capability during transfer. We address this with a stable post-training recipe combining behavior-cloning distillation, critic warm-up, and conservative on-policy updates. To safely exploit the full kinematic dexterity, we introduce a joint-space Geometric Fabric that mediates between the RL policy and the robot. We distill post-trained teachers into perceptive students that zero-shot sim-to-real transfer on two embodiments: a 23 DoF Kuka-Allegro with two RGB cameras, and a 29 DoF Flexiv-Sharpa with two RGB cameras and five vision-based tactile sensors, and can solve long-horizon tasks from challenging initial states with dexterity at human-level speed.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis