Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
AuthorsVarun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev, Animesh Garg
Resources
A small value model helps large robot policies learn from their own mistakes, substantially improving manipulation without more human demonstrations.
Key results
Success after 10 iterations of Q-only self-improvement, up from 93%.
Success after 10 self-improvement iterations, up from 83.8%.
Mean success after self-improvement across the five evaluated benchmarks.
Success after five iterations, up from a 40% frozen-BC baseline.
Success after five iterations, up from 25% without self-improvement.
RoboTwin planning-step latency on an NVIDIA L40S.
What the paper found
Beyond Imitation introduces Q-Planning, a self-improvement method for robot manipulation that keeps a large Behaviour Cloning policy frozen and trains only a smaller off-policy Q-function. The Q-function uses Q-chunking to treat action sequences as super-actions and HL-Gauss categorical regression for stable value learning under sparse terminal rewards, with separate DINOv2 visual and T5 language encoders. At inference, it scores multi-modal three-step flow-matching action chunks sampled from FastWAM, then executes a single softmax Q-weighted average rather than running expensive iterative search. Deployment rollouts, including failures, are added to replay and used for Q-only updates, allowing the system to learn from data that Behaviour Cloning cannot imitate. Across ten self-improvement iterations, LIBERO-10 rises from 93% to 99%, RoboTwin rises from 83.8% to 91.4%, and mean success reaches 97.6%. On real bimanual tasks, five iterations improve stack-cups from 40% to 90% and insert-wallet from 25% to 80%, without human intervention or policy updates; successful-rollout SFT stalls at 55% and 30%. The approximately 1B-parameter Q-function operates in real time at 400 ms per planning step on an NVIDIA L40S, showing a scalable alternative to fine-tuning multi-billion-parameter vision-language-action models.
Original abstract
Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter models underpinning modern robot policies. We propose Q-Planning, which equips a large visuomotor BC policy with a small off-policy Q-function. Because a Q-function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value-guided action selection at inference (a single-step Q-weighted average over BC draws) and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self-improvement lift every benchmark score we tested (LIBERO-10 93% to 99%, RoboTwin 83.8% to 91.4%) and shorten successful episodes on the near-ceiling suites (LIBERO-Object, LIBERO-Goal). On two contact-rich bimanual real-robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack-cups 40% to 90% and insert-wallet 25% to 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q-Planning is the only method, among Best-of-N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.