Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
AuthorsTongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li
Resources
This paper makes robot fine-tuning smarter by turning sparse success/failure signals into more informative, stage-aware training weights, boosting real-world task success on bimanual manipulation.
Key results
baseline success rate improved to 92% with HABC
baseline success rate improved to 88% with HABC
baseline success rate improved to 38% with HABC
best reported final success rate
best reported final success rate
best reported final success rate
What the paper found
Hierarchical Advantage-Weighted Behavior Cloning, or HABC, is a new online RL fine-tuning method for Vision-Language-Action models that targets a core bottleneck in robot post-training: each rollout yields only a single success-or-failure label, yet the actor needs transition-level supervision. The paper, from ACE Robotics, Tsinghua University, and The Chinese University of Hong Kong, shows that a binary episode outcome actually contains two separable signals: viability, meaning whether a state can still lead to success, and efficiency, meaning whether a transition is moving toward faster completion. HABC trains a dual-head critic, with a viability head supervised on all outcome-labeled policy-execution windows and an efficiency head trained only on successful trajectories, then merges their one-step advantages with a state-adaptive gate that emphasizes viability when success is uncertain and shifts toward efficiency once viability is high. It also introduces intervention-aware credit assignment, labeling only the post-intervention policy-execution suffix so human corrections are not incorrectly credited or penalized. On three real-robot bimanual deformable-object tasks with a π0.5 base model, HABC improves success from SFT baselines of 36%, 44%, and 12% to 92%, 88%, and 38%, while also reducing successful trajectory length by 55, 162, and 32 frames, respectively, showing both better recovery and more efficient completion.
Original abstract
When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary outcome (success or failure), yet the actor update requires per-transition supervision. Existing approaches commonly reduce this sparse outcome to a single scalar reward or advantage signal, which conflates distinct forms of transition-level feedback and provides limited guidance once basic task success becomes achievable. First, a single scalar signal conflates the two objectives of viability and efficiency; once basic success is achieved, the binary label provides no gradient to distinguish efficient completions from slow ones. Second, real-world rollouts mix autonomous and intervention segments; naively assigning episode outcomes across these boundaries introduces incorrect credit assignment. To address these issues, we propose Hierarchical Advantage-Weighted Behavior Cloning (HABC), which trains separate critic heads for these two objectives on different data subsets and combines their outputs with a state-adaptive balance. A state-adaptive gate $g_t$ merges their one-step advantages, prioritizing viability when success is uncertain and shifting to efficiency only when viability is high, and converts the result into per-transition weights on the actor loss. Intervention-aware credit assignment further restricts outcome labels to segments executed by the current policy, preventing supervision from leaking across intervention boundaries. In real-robot experiments on three contact-rich bimanual tasks, HABC raises success from supervised fine-tuning (SFT) baselines of 36%, 44%, and 12% to 92%, 88%, and 38%.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.