Robot See Robot Do
RSRD takes in 1) a multi-view object scan and 2) a monocular demonstration video. By creating part-aware 3D representations using GARField (parts, toggle for clusters) and DINOv2 (tracking SE3 pose), these smartphone-captured inputs can generate these 4D reconstructions: These 4D reconstructions are rendered in-browser! If you think that's cool, check out Viser! After recovering 3D part motion, RSRD optimizes grasps and robot motions to reproduce the 4D reconstructions. These can be physically executed on a real robot to produce the demonstrated motion: RSRD's visual imitation is object-centric, allowing it to adapt to different object orientations with the same demo: 0 degrees rotated 180 degrees rotated 0 degrees rotated 30 degrees rotated 45 degrees rotated 4D Differentiable Part Models 4D-DPM decomposes objects into parts with GARField, and trains part-centric feature fields on top of these. Each part is assigned a trainable 6D pose parameter which is optimized with gradient descen
Robot See Robot Do Robot See 👁️ Robot Do 🦾 Imitating Articulated Object Manipulation with Monocular 4D Reconstruction Justin Kerr* Chung Min Kim* Mingxuan Wu Brent Yi Qianqian Wang Ken Goldberg Angjoo Kanazawa UC Berkeley * Denotes Equal Contribution CoRL 2024 (Oral) Paper </Code> Data TL;DR : Robot See Robot Do uses a 4D D ifferentiable P art M odel (4D-DPM) to visually imitate articulated motions from an object scan and single monocular video. Humans imitate manipulation by watching object motion, not hand motion. RSRD does the same. This enables imitation from a single video robust to ori
related reading
- Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wildlift4d.github.io
- Differentiable Robot Renderingdrrobot.cs.columbia.edu
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- State of Robot Learning, December 2025vedder.io
- Humanoid Atlas | Humanoid Robot Supply Chain Map, OEM Database & Industry Analysishumanoids.fyi
- A Unified Model for Motion-Conditioned Robot Co-designtransformer-transformer.github.io
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Macrodata Labs — better data for better robotsmacrodata.co
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- S4Darxiv.org
- eventual.aieventual.ai
- Kimodo: Scaling Controllable Human Motion Generationresearch.nvidia.com