NTH

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

AuthorsTongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li

June 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes robot fine-tuning smarter by turning sparse success/failure signals into more informative, stage-aware training weights, boosting real-world task success on bimanual manipulation.

Key results

36%
Pencil Pouch SFT success

baseline success rate improved to 92% with HABC

44%
Paper Bag SFT success

baseline success rate improved to 88% with HABC

12%
Snack Bag SFT success

baseline success rate improved to 38% with HABC

92%
Pencil Pouch HABC success

best reported final success rate

88%
Paper Bag HABC success

best reported final success rate

38%
Snack Bag HABC success

best reported final success rate

What the paper found

Hierarchical Advantage-Weighted Behavior Cloning, or HABC, is a new online RL fine-tuning method for Vision-Language-Action models that targets a core bottleneck in robot post-training: each rollout yields only a single success-or-failure label, yet the actor needs transition-level supervision. The paper, from ACE Robotics, Tsinghua University, and The Chinese University of Hong Kong, shows that a binary episode outcome actually contains two separable signals: viability, meaning whether a state can still lead to success, and efficiency, meaning whether a transition is moving toward faster completion. HABC trains a dual-head critic, with a viability head supervised on all outcome-labeled policy-execution windows and an efficiency head trained only on successful trajectories, then merges their one-step advantages with a state-adaptive gate that emphasizes viability when success is uncertain and shifts toward efficiency once viability is high. It also introduces intervention-aware credit assignment, labeling only the post-intervention policy-execution suffix so human corrections are not incorrectly credited or penalized. On three real-robot bimanual deformable-object tasks with a π0.5 base model, HABC improves success from SFT baselines of 36%, 44%, and 12% to 92%, 88%, and 38%, while also reducing successful trajectory length by 55, 162, and 32 frames, respectively, showing both better recovery and more efficient completion.

Original abstract

When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary outcome (success or failure), yet the actor update requires per-transition supervision. Existing approaches commonly reduce this sparse outcome to a single scalar reward or advantage signal, which conflates distinct forms of transition-level feedback and provides limited guidance once basic task success becomes achievable. First, a single scalar signal conflates the two objectives of viability and efficiency; once basic success is achieved, the binary label provides no gradient to distinguish efficient completions from slow ones. Second, real-world rollouts mix autonomous and intervention segments; naively assigning episode outcomes across these boundaries introduces incorrect credit assignment. To address these issues, we propose Hierarchical Advantage-Weighted Behavior Cloning (HABC), which trains separate critic heads for these two objectives on different data subsets and combines their outputs with a state-adaptive balance. A state-adaptive gate $g_t$ merges their one-step advantages, prioritizing viability when success is uncertain and shifting to efficiency only when viability is high, and converts the result into per-transition weights on the actor loss. Intervention-aware credit assignment further restricts outcome labels to segments executed by the current policy, preventing supervision from leaking across intervention boundaries. In real-robot experiments on three contact-rich bimanual tasks, HABC raises success from supervised fine-tuning (SFT) baselines of 36%, 44%, and 12% to 92%, 88%, and 38%.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis