[2605.22748] Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
Abstract:Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This failure stems from the dominant single-agent paradigm for physical applications, where other actors are ignored or treated as environmental noise, preventing effective coordination. Here we show that multi-agent reinforcement learning provides the essential safety scaffolding required for real-world interaction. Using high-speed quadrotor racing as a high-stakes testbed, we train agents to navigate complex aerodynamic interactions and strategic maneuvering with a variable number of racers. Through league-based self-play, agents evolve sophisticated anticipatory behaviors, including proactive collision avoidance, overtaking, and handling multi-agent physical interactions, including aerodynamic downwash. Our agents outperform a champion-level human pilot in multi-player races at speeds exceeding 22 m/s, while simultaneously reducing collision rates by 50 % compared to state-of-the-art single-agent baselines. Crucially, training with diverse artificial agents enables zero-shot generalization to safer human interaction. These results suggest that the path to robust robotic co-existence lies not in isolated safety constraints, but in the rigorous demands of multi-agent interaction. Multimedia materials are available at: this https URL
# link_1jp1vpbvi49.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Ismail Geles; Leonard Bauersfeld; Markus Wulfmeier; Davide Scaramuzza - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2605.22748 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2605.22748v2 - Producer
Explore this link on the map →saved by
related reading
- Why multi-agent safety is important — LessWronglesswrong.com
- Towards self-driving codebases · Cursorcursor.com
- Human-compatible driving partners through data-regularized self-play reinforcement learningarxiv.org
- How we built our multi-agent research system \ Anthropicanthropic.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- State of Robot Learning, December 2025vedder.io
- Building reliable sim driving agents by scaling self-playarxiv.org
- Scaling long-running autonomous coding · Cursorcursor.com
- Multi-agent safety — AI Alignment Forumalignmentforum.org
- [2502.14143] Multi-Agent Risks from Advanced AIarxiv.org
- VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémonarxiv.org
- Getting Up to Speed on Multi-Agent Systems, Part 1: The Landscapechristophermeiklejohn.com