Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
Junyi Zhang1 Charles Herrmann2 Junhwa Hur2 Eric Chen3 Varun Jampani4 Deqing Sun2 * Ming-Hsuan Yang2,5 * 1 Shanghai Jiao Tong University 2 Google Research 3 UIUC 4 Stability AI 5 UC Merced (*: equal contribution) CVPR 2024 [Paper (updated)] [Arxiv] [Code (new!)] [BibTeX] While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation of the features of current foundation models under simple post-processing. We show that incorporating this information can markedly enhance semantic correspondence performance with simple but effective solutions in both zero-shot and supervised settings. We also construct a new challenging benchmark for semantic correspondence built from an existing animal pose estimation dataset, for both pre-training validat
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence Junyi Zhang 1 Charles Herrmann 2 Junhwa Hur 2 Eric Chen 3 Varun Jampani 4 Deqing Sun 2 * Ming-Hsuan Yang 2,5 * 1 Shanghai Jiao Tong University 2 Google Research 3 UIUC 4 Stability AI 5 UC Merced (*: equal contribution) CVPR 2024 (a) SD+DINO [1] struggles at “telling left from right” (red solid lines). (b) Our method significantly improves semantic correspondence. (c) Qualitative comparison with state-of-the-art methods in cases with extreme vie
Explore this link on the map →saved by
related reading
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondencesd-complements-dino.github.io
- When Models Manipulate Manifolds: The Geometry of a Counting Tasktransformer-circuits.pub
- PLA: Language-Driven Open-Vocabulary 3D Scene Understandingarxiv.org
- Language-Grounded Indoor 3D Semantic Segmentation in the Wildarxiv.org
- Video models are zero-shot learners and reasonersarxiv.org
- Regularized by Score Matching (LeCun)papers.nips.cc
- [2004.06165] Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasksarxiv.org
- [2008.05711] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3Darxiv.org
- 3D-LFM: Lifting Foundation Model3dlfm.github.io
- [2603.14482] V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learningarxiv.org
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3Dresearch.nvidia.com