2023-ConceptFusion.pdf
concept-fusion.github.io · 7,153 words · saved by 1 readers
N/A
ConceptFusion: Open-set Multimodal 3D Mapping Krishna Murthy Jatavallabhula1 , Alihusein Kuwajerwala2,† , Qiao Gu3,† , Mohd Omama4,† , Tao Chen1 , Alaa Maalouf1 , Shuang Li1 , Ganesh Iyer7,‡ , Soroush Saryazdi8 , Nikhil Keetha5 , Ayush Tewari1 , Joshua B. Tenenbaum1 , Celso Miguel de Melo6 , K. Madhava Krishna4 , Liam Paull2 , Florian Shkurti3 , and Antonio Torralba1 1 MIT, 2 Université de Montréal, 3 University of Toronto, 4 IIIT Hyderabad, 5 CMU, 6 DEVCOM Army Research Lab, 7 Amazon, 8 Concordia…
saved by
related reading
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org
- Multi-View Transformer for 3D Visual Groundingarxiv.org
- ReferIt3D: Neural Listeners for Fine-Grained 3D Object Identification in Real-World Scenesecva.net
- MDETR - Modulated Detection for End-to-End Multi-Modal Understandingarxiv.org
- ReferIt3D Benchmarksreferit3d.github.io
- Language-Grounded Indoor 3D Semantic Segmentation in the Wildarxiv.org
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- ImageBind: Holistic AI learning across six modalitiesai.facebook.com
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com
- [2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigmarxiv.org
- AI 3D Model Generator: Create 3D from Text & Images | Meshymeshy.ai
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org