OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effectively control mobile systems, and grasping models that can handle a wide range of objects. Despite these advancements, general-purpose applications of robotics still lag behind, even though they rely on these fundamental capabilities of recognition, navigation, and grasping. In this paper, we adopt a systems-first approach to develop a new Open Knowledge-based robotics framework called OK-Robot. By combining Vision-Language Models (VLMs) for object detection, navigation primitives for movement, and grasping primitives for object manipulation
OK-Robot : What Really Matters in Integrating Open-Knowledge Models for Robotics Peiqi Liu* 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Yaswanth Orru* 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Jay Vakil 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Chris Paxton 2 2 {}^{2} start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Nur Muhammad Mahi Shafiullah2 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Lerrel Pinto2 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT New York University 1 1 {}^{1} start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT , AI at
Explore this link on the map →related reading
- A VLA with Open-World Generalizationpi.website
- Explore | alphaXivalphaxiv.org
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- State of Robot Learning, December 2025vedder.io
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- The ‘ChatGPT Moment’ in Robotics and beyondparitoshmohan.substack.com
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- A Steerable Model with Emergent Capabilitiespi.website
- Abrar Anwarabraranwar.github.io
- A VLA with Open-World Generalizationphysicalintelligence.company
- [2401.14403] Adaptive Mobile Manipulation for Articulated Objects In the Open Worldar5iv.labs.arxiv.org