DINOv3
DINOv3 scales self-supervised learning (SSL) for images to produce our strongest universal vision backbones, enabling breakthrough performance across diverse domains.
INTRODUCING DINOV3 Self-supervised learning for vision at unprecedented scale DINOv3 scales self-supervised learning (SSL) for images to produce our strongest universal vision backbones, enabling breakthrough performance across diverse domains. DINOV3 OVERVIEW Cutting-edge image representations, trained without human supervision We scaled unsupervised training to 7B-parameter models and 1.7B image datasets, using a fraction of compute compared to weakly-supervised methods. Despite keeping backbones frozen during evaluation, they achieve absolute state-of-the-art performance across…
saved by
related reading
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Self-supervised learning: The dark matter of intelligenceai.facebook.com
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- 2403.09611.pdfarxiv.org
- Big Self-Supervised Models are Strong Semi-Supervised Learnersarxiv.org
- [2603.14482] V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learningarxiv.org
- Video models are zero-shot learners and reasonersarxiv.org
- [2203.09795] Three things everyone should know about Vision Transformersarxiv.org
- [2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigmarxiv.org
- Scaling Video Pretraining with Imagination Modelsinductionlabs.com
- [2103.00020] Learning Transferable Visual Models From Natural Language Supervisionarxiv.org
- Stand-Alone Self-Attention in Vision Models - NeurIPS-2019-stand-alone-self-attention-in-vision-models-Paper.pdfpapers.nips.cc