[2303.15771] TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigation
Abstract:Effective use of camera-based vision systems is essential for robust performance in autonomous off-road driving, particularly in the high-speed regime. Despite success in structured, on-road settings, current end-to-end approaches for scene prediction have yet to be successfully adapted for complex outdoor terrain. To this end, we present TerrainNet, a vision-based terrain perception system for semantic and geometric terrain prediction for aggressive, off-road navigation. The approach relies on several key insights and practical considerations for achieving reliable terrain modeling. The network includes a multi-headed output representation to capture fine- and coarse-grained terrain features necessary for estimating traversability. Accurate depth estimation is achieved using self-supervised depth completion with multi-view RGB and stereo inputs. Requirements for real-time performance and fast inference speeds are met using efficient, learned image feature projections. Furthermore, the model is trained on a large-scale, real-world off-road dataset collected across a variety of diverse outdoor environments. We show how TerrainNet can also be used for costmap prediction and provide a detailed framework for integration into a planning module. We demonstrate the performance of TerrainNet through extensive comparison to current state-of-the-art baselines for camera-only scene prediction. Finally, we showcase the effectiveness of integrating TerrainNet within a complete autonomous-driving stack by conducting a real-world vehicle test in a challenging off-road scenario.
TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigation Effective use of camera-based vision systems is essential for robust performance in autonomous off-road driving, particularly in the high-speed regime. Despite success in structured, on-road settings, current end-to-end approaches for scene prediction have yet to be successfully adapted for complex outdoor terrain. To this end, we present TerrainNet, a vision-based terrain perception system for semantic and geometric terrain prediction for aggressive, off-road navigation. The approach relies on several key insi
Explore this link on the map →saved by
related reading
- roboticsproceedings.org/rss19/p103.pdfroboticsproceedings.org
- [2209.10788] How Does It Feel? Self-Supervised Costmap Learning for Off-Road Vehicle Traversabilityarxiv.org
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org
- Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3Dresearch.nvidia.com
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- When Models Manipulate Manifolds: The Geometry of a Counting Tasktransformer-circuits.pub
- How I built my own Tesla-style self-driving AI | Transformer Lablab.cloud
- [2509.02722] Planning with Reasoning using Vision Language World Modelarxiv.org
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- Multi-View Transformer for 3D Visual Groundingarxiv.org