Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
AuthorsTyler Ga Wei Lum, Kushal Kedia, C. Karen Liu, Jeannette Bohg
Resources
This work shows that teaching a dexterous robot to 'play' first can make it far better at precise real-world assembly tasks like tight insertions and screwing.
Key results
Play2Perfect is reported to be 33x more sample-efficient than RL from scratch.
Scratch with dense reward needed over 100 hours to match the pretrained policy on simplified tight insertion.
Play2Perfect reached the same near-perfect success in 4 hours on simplified tight insertion.
Simulation success rate for Play2Perfect at 4 mm clearance.
Simulation success rate for Play2Perfect at 1 mm clearance.
Zero-shot real-world tight-insertion success at 0.5 mm clearance.
What the paper found
Play2Perfect, from Stanford University and Cornell University, argues that dexterous assembly should be learned by first learning task-agnostic “play” and then specializing with sparse-reward RL. The policy is pretrained with SAPG on procedurally generated cuboids and cylinders, using 6D keypoint-based pose control, random goal trajectories, and a tight 1 cm success tolerance to force in-hand manipulation rather than fixed-grasp transport. Finetuning then maps CAD-defined assembly steps into sparse goal sequences via assembly-by-disassembly, including pre-insertion and threaded screw poses, while reusing the same robot prior. The key finding is that this prior is 33x more sample-efficient than training from scratch, even against a dense, hand-shaped reward baseline that still needed over 100 hours to match the pretrained policy’s 4-hour result on simplified tight insertion. In simulation, Play2Perfect reaches 95% success at 4 mm clearance, 92% at 1 mm, and 80% at 0.2 mm, while the play-only policy collapses near 4 mm. Zero-shot sim-to-real transfer is demonstrated on a 22-DoF Sharpa hand with a 7-DoF KUKA iiwa 14 arm and FoundationPose tracking, achieving 60% success at 0.5 mm clearance on tight insertion and over 50% success on multi-part assembly and screwing.
Original abstract
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection for imitation learning difficult, and sparse-reward, making direct exploration with reinforcement learning (RL) intractable. Consequently, prior work has made progress by structuring the problem with specialized grippers, tool attachments, and environment fixtures. In this work, we argue that before a robot can perfect precise assembly, it must first learn to play. We further ask the question: what factors in the process of learning to play matter for precise assembly? We propose Play2Perfect, an RL framework for task-agnostic pretraining through play on diverse objects and goals, which is then perfected on precise assembly. The goal of play is to acquire reusable manipulation priors, such as grasping, in-hand reorientation and pose reaching. Finetuning then adapts this general prior to assembly, focusing exploration on the final contact-rich, high-precision interactions needed for success. We systematically study key design choices in play pretraining, including object diversity, training objective, trajectory diversity, and goal precision. We show that our prior is 33x more sample-efficient than RL training from scratch, even when provided with dense, multi-stage rewards. We demonstrate zero-shot sim-to-real transfer, achieving 60% success on tight insertions with only 0.5 mm contact clearance, and over 50% success on long-horizon multi-part assembly and screwing.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.