Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this training paradigm has the potential to learn visual representations from diverse sources, ranging from natural to aerial images—using a single algorithm. DINOv3 is a major milestone toward realizing this vision by leveraging simple yet effective strategies. First, it scales both dataset and model size by careful data preparation, design, and optimization. Second, a new method called Gram anchoring effectively addresses the known yet unsolved issue of dense feature maps degrading during long training schedules. Finally, post-hoc strategies are applied that further enhance the models’ flexibility with respect to resolution, model size, and alignment with text. DINOv3 is a versatile vision foundation model that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models.
In this video, I talk about the following: What is new in DINOv3? How is the DinoV3 model trained? What is Gram Anchoring? How does the DinoV3 model perform?
For more details, please look at https://ai.meta.com/dinov3/ and https://ai.meta.com/blog/dinov3-self-... and https://arxiv.org/pdf/2508.10104
Siméoni, Oriane, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov et al. "Dinov3." arXiv preprint arXiv:2508.10104 (2025).
Thanks for watching!
LinkedIn: http://aka.ms/manishgupta
HomePage: https://sites.google.com/view/manishg/