3D-LFM: Lifting Foundation Model
3D pose estimation on random videos from the internet using a single model for all the following categories. We show 3D pose estimation on various deformable categories from the internet, including OpenAI's SORA videos. No camera information is required, allowing 3D-LFM to work out of the box on many of these categories, provided 2D landmarks are available (in any order or joint connectivity). A single model for 30+ object categories. The 3D-LFM scales to multiple categories (30+ in our experiments), managing diverse landmark configurations through proposed architectural changes. See our paper for more details. Key: Red for ground truth, blue for predictions. Fish is completely OOD! The 3D-LFM is capable in recognizing and reconstructing objects in configurations and with a number of landmarks never encountered during its training phase. Above object categories are never seen by the model in training. The lifting of 3D structure and camera from 2D landmarks is at the cornerstone of the
3D-LFM: Lifting Foundation Model 3D-LFM: Lifting Foundation Model Mosam Dabhi 1 , László A. Jeni 1 , Simon Lucey 2 1 Carnegie Mellon University, 2 The University of Adelaide CVPR, 2024 Paper --> arXiv Video Inference Press 🤗 Poster Tweet --> Chat with 3D-LFM GPT Tweet --> 3D-LFM is a universal 2D-3D lifting model capable of handling multiple object categories using a single model. It doesn't assume object knowledge and exploits the permutation equivariance of transformers to learn a category-agnostic lifting model. Moreover, it is an efficient model because it exploits geometric reasoning to
Explore this link on the map →saved by
related reading
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org
- ReferIt3D: Neural Listeners for Fine-Grained 3D Object Identification in Real-World Scenesecva.net
- [2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigmarxiv.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- Multi-View Transformer for 3D Visual Groundingarxiv.org
- 3DMV3dmv2023.github.io
- Video models are zero-shot learners and reasonersarxiv.org
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org
- Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3Dresearch.nvidia.com
- Explore | alphaXivalphaxiv.org
- RTFM: A Real-Time Frame Model | World Labsworldlabs.ai
- ReferIt3D Benchmarksreferit3d.github.io