RT-2: Vision-Language-Action Models
robotics-transformer2.github.io · 1,592 words · saved by 2 readers
Project page for RT-2
RT-2: Vision-Language-Action Models --> --> Your browser does not support the video tag. RT2: Vision-Language-Action Models RT-2 model picking up object given the prompt "pick up the extinct animal." RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control Anthony Brohan Noah Brown Justice Carbajal Yevgen Chebotar Xi Chen Krzysztof Choromanski Tianli Ding Danny Driess Avinava Dubey Chelsea Finn Pete Florence Chuyuan Fu Montse Gonzalez Arenas Keerthana Gopalakrishnan Kehang Han Karol Hausman Alex Herzog Jasmine Hsu Brian Ichter Alex Irpan Nikhil Joshi Ryan Julian Dmitry Kal
saved by
related reading
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- A VLA with Open-World Generalizationpi.website
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- $π_0$: A Vision-Language-Action Flow Model for General Robot Controlalphaxiv.org
- A Steerable Model with Emergent Capabilitiespi.website
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learningalphaxiv.org
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Explore | alphaXivalphaxiv.org
- Physical Intelligence (π)pi.website