Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
AuthorsIsmail Geles, Leonard Bauersfeld, Markus Wulfmeier, Davide Scaramuzza
Resources
By training racing drones against each other in multi-agent self-play, the system learns safer, faster, and more human-compatible maneuvers than single-agent methods.
Key results
Physical races were conducted on a 75-meter circuit with seven gates, including the Split-S maneuver.
The real-world quadrotors used in the experiments were 220-gram racing drones.
The learned policies were evaluated in real-world multi-player races at speeds reaching 22 m/s.
The paper reports real-world racing at accelerations of up to 7 g.
League-play reduced collisions by 50% relative to single-agent baselines in both real-world and large-scale evaluation contexts.
Large-scale simulated evaluation was performed across 64,000 four-player races to compare training paradigms.
What the paper found
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning shows that the main obstacle in physical autonomy is not single-agent control, but learning to coexist with other agents in the same space. The authors train quadrotor racing policies in Flightmare and Agilicious using recurrent PPO with LSTM layers, a Perceiver-based permutation-invariant attention encoder for variable numbers of opponents, league-play self-play, and a particle-based downwash model. In real-world races on a 75-meter Split-S track with 220-gram drones at up to 22 m/s and 7 g, the league-play policy beat five-time Swiss drone racing champion Marvin Schaepper in multi-player settings while reducing collisions by 50% relative to single-agent baselines. In 1v1 races, the policy completed 100% of trials versus the human pilot’s 53.33%; in 3v1 races, it achieved 91.67% completion versus 58.33% for the human. Across 64,000 simulated four-player races, league-play delivered 23.30% crash rate versus 77.99% for single-agent training, and it remained robust up to eight agents, where completion still reached 56%. The key novelty is that safety emerged from diverse opponent exposure rather than explicit hand-coded avoidance: the learned critic function became opponent-aware and anticipatory, enabling overtakes, yielding, and wake avoidance. An ablation without the Perceiver encoder degraded performance, confirming that architectural handling of variable opponent sets is essential for multi-agent racing.
Original abstract
Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This failure stems from the dominant single-agent paradigm for physical applications, where other actors are ignored or treated as environmental noise, preventing effective coordination. Here we show that multi-agent reinforcement learning provides the essential safety scaffolding required for real-world interaction. Using high-speed quadrotor racing as a high-stakes testbed, we train agents to navigate complex aerodynamic interactions and strategic maneuvering with a variable number of racers. Through league-based self-play, agents evolve sophisticated anticipatory behaviors, including proactive collision avoidance, overtaking, and handling multi-agent physical interactions, including aerodynamic downwash. Our agents outperform a champion-level human pilot in multi-player races at speeds exceeding 22 m/s, while simultaneously reducing collision rates by 50 % compared to state-of-the-art single-agent baselines. Crucially, training with diverse artificial agents enables zero-shot generalization to safer human interaction. These results suggest that the path to robust robotic co-existence lies not in isolated safety constraints, but in the rigorous demands of multi-agent interaction. Multimedia materials are available at: https://rpg.ifi.uzh.ch/marl
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.