RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standardization, either by specifying fixed evaluation tasks and environments, or by hosting centralized “robot challenges”, and do not readily scale to evaluating generalist policies across a broad range
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies Pranav Atreya ∗,1 Karl Pertsch ∗,1,2 Tony Lee ∗,2 Moo Jin Kim 2 Arhan Jain 3 Artur Kuramshin 4 Clemens Eppner 5 Cyrus Neary 4 Edward Hu 6 Fabio Ramos 5 Jonathan Tremblay 5 Kanav Arora 3 Kirsty Ellis 4 Luca Macesanu 7 Matthew Leonard 6 Meedeum Cho 8 Ozgur Aslan 4 Shivin Dass 7 Jie Wang 6 Xingfang Yuan 6 Xuning Yang 5 Abhishek Gupta 3 Dinesh Jayaraman 6 Glen Berseth 4 Kostas Daniilidis 6 Roberto Martin-Martin 7 Youngwoon Lee 8 Percy Liang 2 Chelsea Finn 2 Sergey Levine 1 https://robo-arena.github.io Abstract Comprehensive,
Explore this link on the map →saved by
related reading
- What Do Robotics Leaderboards Tell Us About The State of Robot Learning?itcanthink.substack.com
- How do We Quantify Progress in Robotics? - by Chris Paxtonitcanthink.substack.com
- State of Robot Learning, December 2025vedder.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- A VLA with Open-World Generalizationpi.website
- How Can We Make Robotics More like Generative Modeling? | Eric Jangevjang.com
- RLDG: Robotic Generalist Policy Distillation via Reinforcement Learningarxiv.org
- [2503.06814] Acknowledgementsar5iv.labs.arxiv.org
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulationtoyotaresearchinstitute.github.io
- A Steerable Model with Emergent Capabilitiespi.website
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Abrar Anwarabraranwar.github.io