NTH

Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

AuthorsYijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang

August 19, 2026 2 min read
Watch on YouTube
The one-line take

Robo-Dopamine 2.0 gives robot policies a history-aware, OOD-sensitive reward signal to recover from mistakes and improve manipulation reliability.

Key results

0.986
Mean VOC with history panels

Improved from 0.967 using static inputs

0.958
OOD-robust VOC

Improved from 0.906 with history-conditioned panels

400K
Pairwise training budget

Signed pairs used across two curriculum stages

0.9872
Signed-Hop mean VOC

Achieved with 25% replay, versus 0.9858 for shuffled sampling

86.8%
RoboTwin mean success

Full history and OOD reward condition

71
Real-world successful insertions

Successful insertions out of 80 attempts

What the paper found

Robo-Dopamine 2.0 is a process reward model for refining vision-language-action policies when robotic manipulation leaves the demonstration distribution. Instead of judging static before-and-after images, it conditions pairwise progress predictions on ordered temporal context: same-episode expert history during training, observed rollout history online, and source-aligned successful-reference panels for synthetic OOD queries while preserving the edited endpoints. Its OOD-aware signed progress space assigns positive values to valid states, identical values to robustness-preserving changes such as occlusion or distractors, and negative branch values to failures including wrong-object interaction, empty grasps, and non-release, while also representing recovery. The Signed-Hop Curriculum trains coarse global ordering before fine calibration and replays 25% large-transition pairs. Using Qwen3-VL-8B, RoboCasa, LIBERO, and AgiBot World data, history panels raise mean VOC from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958; under a fixed 400K pairwise budget, curriculum training reaches 0.9872 mean VOC versus 0.9858 for shuffled sampling. Frozen reward-model potentials then shape downstream RL: the full condition achieves 86.8% mean RoboTwin success and 71 successful insertions out of 80 real-world attempts. The model also records 96.1% task-completion classification accuracy, exceeding GPT-5.5, Claude Opus-4.8, and Gemini-3-Pro on that evaluation, although this discrete metric is separate from VOC.

Original abstract

Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, while engineered dense rewards are costly and task-specific. Existing learned visual reward models often rely on static before-after observations, causing temporal ambiguity and weak discrimination between robustness-preserving variations and task-invalid failures under out-of-distribution (OOD) execution. We introduce Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface. It combines (1) history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried endpoints, and (2) an OOD-aware signed progress space that represents valid progress, robustness, failure, and recovery. A Signed-Hop Curriculum with transition-aware replay learns coarse execution ordering before fine-grained progress calibration. We also construct an OOD trajectory dataset and a five-family benchmark. Reference panels improve mean visual order consistency (VOC) from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958. With the same 400K pairwise-reward budget, Signed-Hop training with 25% replay reaches 0.9872 mean VOC, compared with 0.9858 for a matched-pool shuffled control. In downstream reinforcement learning, the full model achieves 86.8% mean RoboTwin success and 71/80 successful real-world insertions.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis