Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
Junyi Zhang1 Charles Herrmann2 Junhwa Hur2 Eric Chen3 Varun Jampani4 Deqing Sun2 * Ming-Hsuan Yang2,5 * 1 Shanghai Jiao Tong University 2 Google Research 3 UIUC 4 Stability AI 5 UC Merced (*: equal contribution) CVPR 2024 [Paper (updated)] [Arxiv] [Code (new!)] [BibTeX] While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation of the features of current foundation models under simple post-processing. We show that incorporating this information can markedly enhance semantic correspondence performance with simple but effective solutions in both zero-shot and supervised settings. We also construct a new challenging benchmark for semantic correspondence built from an existing animal pose estimation dataset, for both pre-training validat
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence Junyi Zhang 1 Charles Herrmann 2 Junhwa Hur 2 Eric Chen 3 Varun Jampani 4 Deqing Sun 2 * Ming-Hsuan Yang 2,5 * 1 Shanghai Jiao Tong University 2 Google Research 3 UIUC 4 Stability AI 5 UC Merced (*: equal contribution) CVPR 2024 (a) SD+DINO [1] struggles at “telling left from right” (red solid lines). (b) Our method significantly improves semantic correspondence. (c) Qualitative comparison with state-of-the-art methods in cases with extreme vie
saved by
related reading
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondencesd-complements-dino.github.io
- When Models Manipulate Manifolds: The Geometry of a Counting Tasktransformer-circuits.pub
- Language-Grounded Indoor 3D Semantic Segmentation in the Wildarxiv.org
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org
- Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusionarxiv.org
- [2004.06165] Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasksarxiv.org
- [2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigmarxiv.org
- Video models are zero-shot learners and reasonersarxiv.org
- DINOv3ai.meta.com
- DUSt3R: Geometric 3D Vision Made Easyarxiv.org
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org
- Blueprint-Bench: Testing spatial intelligence in AI models | Andon Labsandonlabs.com