Aberration-Aware Depth-from-Focus
Computer vision methods for depth estimation usually use simple camera models with idealized optics. For modern machine learning approaches, this creates an issue when attempting to train deep networks with simulated data, especially for focus-sensitive tasks like Depth-from-Focus. In this work, we investigate the domain gap caused by off-axis aberrations that will affect the decision of the best-focused frame in a focal stack. We then explore bridging this domain gap through aberration-aware training (AAT). Our approach involves a lightweight network that models lens aberrations at different positions and focus distances, which is then integrated into the conventional network training pipeline. We evaluate the generality of pretrained models on both synthetic and real-world data. Our experimental results demonstrate that the proposed AAT scheme can improve depth estimation accuracy without fine-tuning the model or modifying the network architecture.
Computer vision methods for depth estimation usually use simple camera models with idealized optics. For modern machine learning approaches, this creates an issue when attempting to train deep networks with simulated data, especially for focus-sensitive tasks like Depth-from-Focus. In this work, we investigate the domain gap caused by off-axis aberrations that will affect the decision of the best-focused frame in a focal stack. We then explore bridging this domain gap through aberration-aware training (AAT). Our approach involves a lightweight network that models lens aberrations at different
Explore this link on the map →related reading
- The Little Book of Deep Learningfleuret.org
- Reproducing DeepTFUS | projectsmasonjwang.com
- The Decade of Deep Learning | Leo Gaobmk.sh
- Extending Stein's unbiased risk estimator to train deep denoisers with correlated pairs of noisy imagesproceedings.neurips.cc
- Foundations of Computer Visionvisionbook.mit.edu
- PS$^2$F: Polarized Spiral Point Spread Function for Single-Shot 3D Sensingarxiv.org
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org
- [2602.05970] Inverse Depth Scaling From Most Layers Being Similararxiv.org
- Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3Dresearch.nvidia.com
- Unsupervised Learning with SURE.pdfarxiv.org
- [2312.03079] LooseControl: Lifting ControlNet for Generalized Depth Conditioningarxiv.org
- [2303.15771] TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigationarxiv.org