[2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm --> Computer Science > Computer Vision and Pattern Recognition arXiv:2310.08586 (cs) [Submitted on 12 Oct 2023 ( v1 ), last revised 15 Apr 2025 (this version, v4)] Title: PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm Authors: Haoyi Zhu , Honghui Yang , Xiaoyang Wu , Di Huang , Sha Zhang , Xianglong He , Hengshuang Zhao , Chunhua Shen , Yu Qiao , Tong He , Wanli Ouyang View a PDF of the paper titled PonderV2: Pave the Way for 3D Foundation Model with A Unive
saved by
related reading
- 3D-LFM: Lifting Foundation Model3dlfm.github.io
- Learning 3D object-centric representation through predictionarxiv.org
- 2D Amodal Instance Segmentation Guided by 3D Shape Priorecva.net
- Enhancing 3D Object Detection with 2D Detection-Guided Query Anchorsarxiv.org
- Amodal Detection of 3D Objects: Inferring 3D Bounding Boxes from 2D Ones in RGB-Depth Imagescis.temple.edu
- Superquadrics Revisited: Learning 3D Shape Parsing beyond Cuboidsarxiv.org
- S4Darxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Language-Grounded Indoor 3D Semantic Segmentation in the Wildarxiv.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- 3DMV3dmv2023.github.io
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org