NTH

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

AuthorsVarun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev, Animesh Garg

August 27, 2026 2 min read
Watch on YouTube
The one-line take

A small value model helps large robot policies learn from their own mistakes, substantially improving manipulation without more human demonstrations.

Key results

99%
LIBERO-10 final success

Success after 10 iterations of Q-only self-improvement, up from 93%.

91.4%
RoboTwin final success

Success after 10 self-improvement iterations, up from 83.8%.

97.6%
Mean benchmark success

Mean success after self-improvement across the five evaluated benchmarks.

90%
Stack-cups real-robot success

Success after five iterations, up from a 40% frozen-BC baseline.

80%
Insert-wallet real-robot success

Success after five iterations, up from 25% without self-improvement.

400 ms
Planning latency

RoboTwin planning-step latency on an NVIDIA L40S.

What the paper found

Beyond Imitation introduces Q-Planning, a self-improvement method for robot manipulation that keeps a large Behaviour Cloning policy frozen and trains only a smaller off-policy Q-function. The Q-function uses Q-chunking to treat action sequences as super-actions and HL-Gauss categorical regression for stable value learning under sparse terminal rewards, with separate DINOv2 visual and T5 language encoders. At inference, it scores multi-modal three-step flow-matching action chunks sampled from FastWAM, then executes a single softmax Q-weighted average rather than running expensive iterative search. Deployment rollouts, including failures, are added to replay and used for Q-only updates, allowing the system to learn from data that Behaviour Cloning cannot imitate. Across ten self-improvement iterations, LIBERO-10 rises from 93% to 99%, RoboTwin rises from 83.8% to 91.4%, and mean success reaches 97.6%. On real bimanual tasks, five iterations improve stack-cups from 40% to 90% and insert-wallet from 25% to 80%, without human intervention or policy updates; successful-rollout SFT stalls at 55% and 30%. The approximately 1B-parameter Q-function operates in real time at 400 ms per planning step on an NVIDIA L40S, showing a scalable alternative to fine-tuning multi-billion-parameter vision-language-action models.

Original abstract

Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter models underpinning modern robot policies. We propose Q-Planning, which equips a large visuomotor BC policy with a small off-policy Q-function. Because a Q-function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value-guided action selection at inference (a single-step Q-weighted average over BC draws) and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self-improvement lift every benchmark score we tested (LIBERO-10 93% to 99%, RoboTwin 83.8% to 91.4%) and shorten successful episodes on the near-ceiling suites (LIBERO-Object, LIBERO-Goal). On two contact-rich bimanual real-robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack-cups 40% to 90% and insert-wallet 25% to 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q-Planning is the only method, among Best-of-N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →