LIMBO: Learning and Internalizing Model-Free Barrier Objectives for Agile and Safe Whole-Body Control
AuthorsJake Gonzales, Arturo Flores Alvarez, Yu-Ming Chen, Aaron D. Ames, Lillian J. Ratliff, Manikantan Nambi
Resources
LIMBO teaches humanoid robots to discover and internalize safety boundaries so they can perform agile maneuvers without relying on an online safety filter.
Key results
Residual Q-CBF synthesis and task-policy control used 29 dimensions on the Unitree G1 humanoid.
Training used 4096 parallel MuJoCo environments.
QCBF-PD achieved 1.48% hits versus 3.58% for CBF-RL.
QCBF-PD achieved 38.11% hits versus 59.50% for CBF-RL.
QCBF-PD’s median drift was 0.162 m versus 0.944 m for CBF-RL.
Clearance remained between 85.7% and 91.2% as boundary-sampling concentration changed.
What the paper found
LIMBO introduces a method for safe, agile whole-body control that first learns a state–action safety certificate, or Q-CBF, from black-box transitions and a state-based failure specification, then distills that certificate into a task policy. Its residual formulation learns safety in the same 29-dimensional control space used by the task policy, while a frozen base controller preserves nominal balance or locomotion. Risk-guided replay concentrates training near the estimated boundary of recoverability, allowing the same safety specification to produce either crouching or a backward-leaning limbo maneuver. During PPO training, the learned Q-CBF acts as a counterfactual teacher, supplying violation penalties and minimum-deviation action corrections; the deployed policy therefore needs no online safety filter. Experiments used a Unitree G1 humanoid, 4096 parallel MuJoCo environments, 50 Hz control, AMP motion priors, and Kimodo-generated reference motions. In simulated dodgeball, LIMBO’s QCBF-PD policy achieved a 1.48% in-distribution hit rate versus 3.58% for analytical CBF-RL, and 38.11% versus 59.50% on out-of-distribution targets, with zero falls in both settings. On hardware, it transferred directly from simulation without adaptation or filtering; synthetic throws produced 0.162 m median drift versus 0.944 m for CBF-RL. For locomotion beneath a 1.234 m obstacle, boundary sampling shifted the strategy from pelvis dropping to torso leaning, while clearance remained 85.7% to 91.2%, demonstrating behavior discovery through learned safety rather than hand-designed motion targets.
Original abstract
Safe whole-body control requires coordinating collision avoidance and balance under high-dimensional, nonlinear dynamics--making safety certificates difficult to design and reuse across behaviors. We present LIMBO, a framework for synthesizing a state-action control barrier function and distilling its safety structure into a task policy. LIMBO learns the safety certificate from black-box transitions and a state-based failure specification over residual actions around a frozen base controller, making Q-CBF synthesis tractable in the full control dimension while placing the certificate in the task policy's control space. During synthesis, the learned safety value drives risk-guided sampling near the estimated boundary of recoverability; during task learning, it serves as a teacher that provides action-level safety feedback, yielding a robust task policy and alleviating the need for an online safety filter at deployment. We demonstrate LIMBO on a 29-degree-of-freedom humanoid performing dodgeball avoidance and locomotion beneath low obstacles. Beyond scaling learned Q-CBFs to whole-body control, we show that risk-guided boundary sampling provides a theoretically grounded way to explore the edge of recoverability. Under the same safety specification, ceteris paribus, varying the sampling concentration produces strategies ranging from crouching to a novel backward-leaning limbo maneuver. In both settings, the learned policies transfer to hardware without online safety filtering, showing that learned safety synthesis scales to agile whole-body control.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.