Human-compatible driving partners through data-regularized self-play reinforcement learning
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. A central challenge for autonomous vehicles is coordinating with humans. Therefore, incorporating realistic human agents is essential for scalable training and evaluation of autonomous driving systems in simulation. Simulation agents are typically developed by imitating large-scale, high-quality datasets of human driving. However, pure imitation learning agents empirically have high collision rates when executed in a multi-agent closed-loop setting. To build agents that are realistic and effective in closed-loop settings, we propose Human-Regularized PPO (HR-PPO), a multi-agent algorithm where agents are trained through self-play with a small penalty for deviating from a human reference policy. In contrast to prior work, our approach is RL-first and only
Human-compatible driving partners through data-regularized self-play reinforcement learning Daphne Cornelisse New York University cornelisse.daphne@nyu.edu &Eugene Vinitsky New York University eugenevinitsky@nyu.edu Abstract A central challenge for autonomous vehicles is coordinating with humans. Therefore, incorporating realistic human agents is essential for scalable training and evaluation of autonomous driving systems in simulation. Simulation agents are typically developed by imitating large-scale, high-quality datasets of human driving. However, pure imitation learning agents empirically
Explore this link on the map →related reading
- Building reliable sim driving agents by scaling self-playarxiv.org
- [2605.22748] Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learningarxiv.org
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- State of Robot Learning, December 2025vedder.io
- [2312.15122] Scaling Is All You Need: Autonomous Driving with JAX-Accelerated Reinforcement Learningarxiv.org
- pistar06.pdfpi.website
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- So You Think You Can Scale Up Autonomous Robot Data Collection?arxiv.org
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- π*0.6: a VLA That Learns From Experiencephysicalintelligence.company
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com