NTH

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

AuthorsGiovanbattista Gravina, Luca Rossini, Carlo Rizzardo, Arturo Laurenzi, Nikos Tsagarakis

August 10, 2026 2 min read
Watch on YouTube
The one-line take

A reinforcement-learning controller helps a large quadruped keep walking after actuator failures by learning when to change its gait timing.

Key results

68
Robot mass

Mass in kilograms of the Kyon quadruped used for simulation and real-world validation.

8192
Parallel training agents

Number of simulated agents trained in parallel.

50
Control frequency

Deployment policy control rate in hertz.

3
Observation history

Selected proprioceptive history length for the full method, written as H=3.

13°
Unseen ramp generalization

Maximum ramp gradient successfully traversed in sim-to-sim tests.

10
Unseen stair step height

Novel stair step height in centimeters used for sim-to-sim evaluation.

What the paper found

This paper presents a deep reinforcement learning controller for quadruped locomotion after sudden actuator power loss, targeting the limited actuation margins of heavier robots. The policy uses Proximal Policy Optimization with an asymmetric actor–critic architecture: during training, the critic receives privileged fault information, while the deployable actor reconstructs it from proprioceptive observation history through a latent-alignment Mean Squared Error loss. Its key innovation is adding a learnable gait-frequency action alongside joint-position targets, allowing step timing and contact scheduling to adapt to terrain and damage without predefined faulty-leg behaviors. Trained in MuJoCo XLA with 8192 parallel agents, the controller runs at 50 Hz on a 68 kg, 12-joint Kyon quadruped and learns to compensate for randomly timed single-joint torque failures on stepped terrain. Ablations show that increasing history from one observation to two provides the largest gain, while the selected H=3 offers a practical capacity–performance trade-off. In sim-to-sim tests, the policy generalized to previously unseen stairs with 10 cm steps and 0.7 m widths, plus ramps up to 13°, and it transferred zero-shot to the physical robot on flat ground. Training used an NVIDIA GeForce RTX 5090, while the reported hardware validation demonstrates stable locomotion under a sudden knee-joint power loss.

Original abstract

Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis