Robot See Robot Do
RSRD takes in 1) a multi-view object scan and 2) a monocular demonstration video. By creating part-aware 3D representations using GARField (parts, toggle for clusters) and DINOv2 (tracking SE3 pose), these smartphone-captured inputs can generate these 4D reconstructions: These 4D reconstructions are rendered in-browser! If you think that's cool, check out Viser! After recovering 3D part motion, RSRD optimizes grasps and robot motions to reproduce the 4D reconstructions. These can be physically executed on a real robot to produce the demonstrated motion: RSRD's visual imitation is object-centric, allowing it to adapt to different object orientations with the same demo: 0 degrees rotated 180 degrees rotated 0 degrees rotated 30 degrees rotated 45 degrees rotated 4D Differentiable Part Models 4D-DPM decomposes objects into parts with GARField, and trains part-centric feature fields on top of these. Each part is assigned a trainable 6D pose parameter which is optimized with gradient descen
Robot See Robot Do Robot See 👁️ Robot Do 🦾 Imitating Articulated Object Manipulation with Monocular 4D Reconstruction Justin Kerr* Chung Min Kim* Mingxuan Wu Brent Yi Qianqian Wang Ken Goldberg Angjoo Kanazawa UC Berkeley * Denotes Equal Contribution CoRL 2024 (Oral) Paper </Code> Data TL;DR : Robot See Robot Do uses a 4D D ifferentiable P art M odel (4D-DPM) to visually imitate articulated motions from an object scan and single monocular video. Humans imitate manipulation by watching object motion, not hand motion. RSRD does the same. This enables imitation from a single video robust to ori
Explore this link on the map →related reading
- State of Robot Learning, December 2025vedder.io
- Humanoid Atlas | Humanoid Robot Supply Chain Map, OEM Database & Industry Analysishumanoids.fyi
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- Learning dexterity | OpenAIopenai.com
- ReferIt3D: Neural Listeners for Fine-Grained 3D Object Identification in Real-World Scenesecva.net
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- Building a robotics research setup that lives next to my desk – dfdx labsdfdxlabs.com
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- 2412.02676arxiv.org
- Ch. 1 - Introductionmanipulation.csail.mit.edu
- MolmoAct Action Reasoning Models that can Reason in Spacearxiv.org