Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D
The goal of perception for autonomous vehicles is to extract semantic representations from multiple sensors and fuse these representations into a single “bird’s-eye-view” coordinate frame for consumption by motion planning. We propose a new end-to-end architecture that directly extracts a bird’s-eye-view representation of a scene given image data from an arbitrary number of cameras. The core idea behind our approach is to “lift” each image individually into a frustum of features for each camera, then “splat” all frustums into a rasterized bird’s-eye- view grid. By training on the entire camera rig, we provide evidence that our model is able to learn not only how to represent images but how to fuse predictions from all cameras into a single cohesive representation of the scene while being robust to calibration error. On standard bird’s- eye-view tasks such as object segmentation and map segmentation, our model outperforms all baselines and prior work. In pursuit of the goal of learning
Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D Toronto AI Lab Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D Jonah Philion , Sanja Fidler NVIDIA, Vector Institute, University of Toronto ECCV 2020 --> --> The goal of perception for autonomous vehicles is to extract semantic representations from multiple sensors and fuse these representations into a single “bird’s-eye-view” coordinate frame for consumption by motion planning. We propose a new end-to-end architecture that directly extracts a bird’s-e
Explore this link on the map →saved by
related reading
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org
- roboticsproceedings.org/rss19/p103.pdfroboticsproceedings.org
- [2303.15771] TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigationarxiv.org
- Event Cameras - Everything you need to know | by Vikram Setty | Mediummedium.com
- The Annotated JEPA | Elements of a Vector Spaceelonlit.com
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- 3DMV3dmv2023.github.io
- Multi-View Transformer for 3D Visual Groundingarxiv.org
- 3D-LFM: Lifting Foundation Model3dlfm.github.io
- When Models Manipulate Manifolds: The Geometry of a Counting Tasktransformer-circuits.pub
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io