DINO: Self-Supervised Vision Transformers

Опубликовано: 15 Февраль 2026
на канале: Soroush Mehraban
9,042
277

DINO, a remarkable self-supervised method, employs two distinct augmented views of an image to acquire the ability to concentrate on objects within the image and generate distinguishable representations for various image categories. It has outperformed prior self-supervised techniques across a range of vision tasks and impressively achieved an 80.1% accuracy on ImageNet, all while utilizing the ViT-B as its backbone.

Paper link: https://arxiv.org/abs/2104.14294

Table of Content:
00:00 Introduction
03:45 Knowledge Distillation
05:13 DINO
07:40 Multi-crop training
12:31 Avoiding Collapse
16:06 Results


Icon made by Freepik from flaticon.com