Learning Fault-Tolerant Locomotion with Adaptive Gait Timing
AuthorsGiovanbattista Gravina, Luca Rossini, Carlo Rizzardo, Arturo Laurenzi, Nikos Tsagarakis
Resources
A reinforcement-learning controller helps a large quadruped keep walking after actuator failures by learning when to change its gait timing.
Key results
Mass in kilograms of the Kyon quadruped used for simulation and real-world validation.
Number of simulated agents trained in parallel.
Deployment policy control rate in hertz.
Selected proprioceptive history length for the full method, written as H=3.
Maximum ramp gradient successfully traversed in sim-to-sim tests.
Novel stair step height in centimeters used for sim-to-sim evaluation.
What the paper found
This paper presents a deep reinforcement learning controller for quadruped locomotion after sudden actuator power loss, targeting the limited actuation margins of heavier robots. The policy uses Proximal Policy Optimization with an asymmetric actor–critic architecture: during training, the critic receives privileged fault information, while the deployable actor reconstructs it from proprioceptive observation history through a latent-alignment Mean Squared Error loss. Its key innovation is adding a learnable gait-frequency action alongside joint-position targets, allowing step timing and contact scheduling to adapt to terrain and damage without predefined faulty-leg behaviors. Trained in MuJoCo XLA with 8192 parallel agents, the controller runs at 50 Hz on a 68 kg, 12-joint Kyon quadruped and learns to compensate for randomly timed single-joint torque failures on stepped terrain. Ablations show that increasing history from one observation to two provides the largest gain, while the selected H=3 offers a practical capacity–performance trade-off. In sim-to-sim tests, the policy generalized to previously unseen stairs with 10 cm steps and 0.7 m widths, plus ramps up to 13°, and it transferred zero-shot to the physical robot on flat ground. Training used an NVIDIA GeForce RTX 5090, while the reported hardware validation demonstrates stable locomotion under a sudden knee-joint power loss.
Original abstract
Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.