RoboCasa | robocasa-web
RoboCasa is a large-scale simulation framework for training generally capable robots to perform everyday tasks. It features realistic and diverse human-centered environments with a focus on kitchen scenes. We create these environments with the aid of generative AI tools, such as large language models (LLMs) and text-to-image/3D generative models. We provide over 2,500 3D assets across 150+ object categories and dozens of interactable furniture and appliances. As part of the first release, we include a suite of 100 tasks, representing a wide spectrum of everyday activities. Together with the simulated tasks, we offer a dataset of high-quality human demonstrations and leverage automated trajectory generation techniques to significantly expand the amount of training data with little additional cost. In this initial release, we focus on kitchen scenes. To capture the complexity and diversity of real-world environments, we consult numerous architecture and home design magazines and compile
RoboCasa Updates [7/7/2026] Our target composite task datasets have been updated to include per-frame subtask annotations . Every timestep is labeled with a subtask index, atomic-skill name, stage (i.e. pick / place / navigate), and a natural-language instruction, to support hierarchical policy learning. [5/12/2026] v1.0.1: Updated horizon lengths (1.5x increase) across all tasks for consistency. Please update to the latest version for running evals. [2/18/2026] v1.0: RoboCasa365 release, with 365 tasks, 2500+ kitchen scenes, 2200+ hours of robot demonstration data, and benchmarking support. R
Explore this link on the map →related reading
- A VLA with Open-World Generalizationpi.website
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- A Steerable Model with Emergent Capabilitiespi.website
- State of Robot Learning, December 2025vedder.io
- A VLA with Open-World Generalizationphysicalintelligence.company
- Humanoid Atlas | Humanoid Robot Supply Chain Map, OEM Database & Industry Analysishumanoids.fyi
- Generalist - GEN-1: Scaling Embodied Foundation Models to Masterygeneralistai.com
- Learning Novel Skills from Language-Generated Demonstrationsarxiv.org
- Isaac Sim - Robotics Simulation and Synthetic Data Generation | NVIDIA Developerdeveloper.nvidia.com
- Introducing Gemini Robotics and Gemini Robotics-ER, AI models designed for robots to understand, act and react to the physical world. — Google DeepMinddeepmind.google
- Open X-Embodiment: Robotic Learning Datasets and RT-X Modelsarxiv.org
- Language Models can Solve Computer Tasksarxiv.org