Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
AuthorsMariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
AffiliationsSiemens AG, Research and Predevelopment, Germany · Technical University of Munich, Germany
Resources
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Key results
Demonstrations used to initialize Res-HIL across the five tasks
Real-world contact-rich manipulation tasks
Res-HIL success rate during online training
Res-HIL success rate during online training
Average across all five tasks
Ablation result on easy peg-in-hole within the training budget
What the paper found
Res-HIL is a human-in-the-loop residual reinforcement learning framework for adapting dexterous robot skills without relearning the entire policy. It freezes a behavior-cloning base policy and trains a residual actor with TD3 to add corrective actions; the residual output is initialized to zero, preserving the pretrained behavior at the start. Each human intervention supplies two signals: direct behavior-cloning supervision for the residual, computed as the teleoperated action minus the base action, and intervention-aware reward shaping that penalizes autonomous transitions preceding failure. The system uses a UR5 robot, two wrist cameras or an additional global camera, and ResNet-10 visual features pretrained on ImageNet. Across five real-world contact-rich tasks—easy and hard peg insertion, vent insertion, and cable routing with two- or three-camera views—Res-HIL starts from only 20 initial demonstrations. After 10 minutes of online training, it reaches 100% success on easy peg-in-hole and 92% on three-camera cable manipulation, outperforming HIL-SERL, residual fine-tuning without human guidance, and a naive residual baseline. At final evaluation, it reaches at least 88% on every task and 100% on three tasks, while averaging a 5.46 s successful-episode cycle time. Ablations show that zero initialization reduces convergence time from 16 minutes to 8 minutes, reward shaping is necessary for efficient convergence, and removing the intervention-supervised behavior-cloning loss reduces success to 20%, demonstrating that targeted human corrections are more valuable than continuous teleoperation.
Original abstract
Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville
NGU helps RL-trained LLMs stop over-practicing easy problems and spend more effort solving the hard ones.