NTH

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

AuthorsKinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi, Jaegul Choo, Yasuyuki Matsushita

July 4, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how a small simulation-trained corrective policy can make vision-language-action robot policies much more reliable in the real world without extra robot training.

Key results

42%
Real robot success rate baseline

Average success rate of the base VLA on the real FR3 robot

76%
Real robot success rate with residual

Average success rate after zero-shot sim-trained residual RL

5
Task count

Number of manipulation tasks evaluated on simulation and real robot

5 mm
Pose noise max position

Maximum training perturbation for object pose position

0.1
Pose noise max orientation

Maximum training perturbation for object pose orientation in radians

What the paper found

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement, developed by KAIST and Microsoft Research Asia with The University of Tokyo, shows that a frozen Vision-Language-Action base model can be made substantially more robust by a lightweight residual policy trained only in simulation. The core idea is to avoid the usual sim-to-real trap by letting the residual observe only 6-DoF object poses, proprioception, and the base VLA action, then aligning simulation and reality through teleoperation replay so the sim VLA and real VLA share the same action distribution. The residual is optimized with TD3, augmented by hierarchical pose noise up to 5 mm and 0.1 rad plus 0.1 pose dropout, and deployed with FoundationPose and SAM2 confidence-gated tracking. On five tabletop manipulation tasks on a Franka Research 3 robot—Cube Lift, Pick-and-Place, Stack Cube, Close Drawer, and Stand Cup Up—the method transfers zero-shot and raises average real-robot success from 42% to 76%, while also reducing successful episode length by 9–22%. The paper further shows the residual generalizes to a stronger backbone, π0.5, and that residual-corrected rollouts can be reused to retrain the base VLA, creating a self-improvement loop without additional teleoperation.

Original abstract

Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions due to compounding execution errors; Can a reinforcement learning policy trained purely in simulation improve the robustness of real-world VLAs zero-shot? Residual RL, which learns a corrective policy on top of a frozen VLA, offers a natural framework, but existing approaches face a fundamental sim-to-real dilemma: privileged-state methods require lossy distillation for deployment; image-based methods suffer from the visual domain gap; and real-world RL is costly and unsafe. We propose an object-centric residual RL framework that refines VLA actions using object poses, enabling a compact observation space that transfers consistently between simulation and reality. To align the two domains, we additionally replay the same teleoperation demonstrations in simulation to train a sim counterpart of the real-world VLA. The residual RL policy is trained only in simulation with pose noise injection and dropout, and transfers zero-shot to the real robot. Across five manipulation tasks on a real Franka Research 3 (FR3) robot, our method improves the success rate from 42% to 76% zero-shot, and the improved rollouts can be further reused to retrain the base VLA for self-improvement without additional teleoperation. Project page: https://www.microsoft.com/en-us/research/articles/object-centric-residual-rl/

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis