NTH

LIMBO: Learning and Internalizing Model-Free Barrier Objectives for Agile and Safe Whole-Body Control

AuthorsJake Gonzales, Arturo Flores Alvarez, Yu-Ming Chen, Aaron D. Ames, Lillian J. Ratliff, Manikantan Nambi

September 21, 2026 2 min read
Watch on YouTube
The one-line take

LIMBO teaches humanoid robots to discover and internalize safety boundaries so they can perform agile maneuvers without relying on an online safety filter.

Key results

29
Whole-body control dimension

Residual Q-CBF synthesis and task-policy control used 29 dimensions on the Unitree G1 humanoid.

4096
Parallel simulation environments

Training used 4096 parallel MuJoCo environments.

1.48%
In-distribution dodgeball hit rate

QCBF-PD achieved 1.48% hits versus 3.58% for CBF-RL.

38.11%
Out-of-distribution dodgeball hit rate

QCBF-PD achieved 38.11% hits versus 59.50% for CBF-RL.

0.162
Synthetic hardware median drift

QCBF-PD’s median drift was 0.162 m versus 0.944 m for CBF-RL.

85.7%
Low-obstacle clearance range

Clearance remained between 85.7% and 91.2% as boundary-sampling concentration changed.

What the paper found

LIMBO introduces a method for safe, agile whole-body control that first learns a state–action safety certificate, or Q-CBF, from black-box transitions and a state-based failure specification, then distills that certificate into a task policy. Its residual formulation learns safety in the same 29-dimensional control space used by the task policy, while a frozen base controller preserves nominal balance or locomotion. Risk-guided replay concentrates training near the estimated boundary of recoverability, allowing the same safety specification to produce either crouching or a backward-leaning limbo maneuver. During PPO training, the learned Q-CBF acts as a counterfactual teacher, supplying violation penalties and minimum-deviation action corrections; the deployed policy therefore needs no online safety filter. Experiments used a Unitree G1 humanoid, 4096 parallel MuJoCo environments, 50 Hz control, AMP motion priors, and Kimodo-generated reference motions. In simulated dodgeball, LIMBO’s QCBF-PD policy achieved a 1.48% in-distribution hit rate versus 3.58% for analytical CBF-RL, and 38.11% versus 59.50% on out-of-distribution targets, with zero falls in both settings. On hardware, it transferred directly from simulation without adaptation or filtering; synthetic throws produced 0.162 m median drift versus 0.944 m for CBF-RL. For locomotion beneath a 1.234 m obstacle, boundary sampling shifted the strategy from pelvis dropping to torso leaning, while clearance remained 85.7% to 91.2%, demonstrating behavior discovery through learned safety rather than hand-designed motion targets.

Original abstract

Safe whole-body control requires coordinating collision avoidance and balance under high-dimensional, nonlinear dynamics--making safety certificates difficult to design and reuse across behaviors. We present LIMBO, a framework for synthesizing a state-action control barrier function and distilling its safety structure into a task policy. LIMBO learns the safety certificate from black-box transitions and a state-based failure specification over residual actions around a frozen base controller, making Q-CBF synthesis tractable in the full control dimension while placing the certificate in the task policy's control space. During synthesis, the learned safety value drives risk-guided sampling near the estimated boundary of recoverability; during task learning, it serves as a teacher that provides action-level safety feedback, yielding a robust task policy and alleviating the need for an online safety filter at deployment. We demonstrate LIMBO on a 29-degree-of-freedom humanoid performing dodgeball avoidance and locomotion beneath low obstacles. Beyond scaling learned Q-CBFs to whole-body control, we show that risk-guided boundary sampling provides a theoretically grounded way to explore the edge of recoverability. Under the same safety specification, ceteris paribus, varying the sampling concentration produces strategies ranging from crouching to a novel backward-leaning limbo maneuver. In both settings, the learned policies transfer to hardware without online safety filtering, showing that learned safety synthesis scales to agile whole-body control.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis