4 тысяч подписчиков
33 видео
R-CNN: Clearly EXPLAINED!
TrackFormer: Multi-Object Tracking with Transformers
Graph Convolutional Networks (GCN): From CNN point of view
MoCo (+ v2): Unsupervised learning in computer vision
Receptive Fields: Why 3x3 conv layer is the best?
ViTPose: 2D Human Pose Estimation
Faster R-CNN: Faster than Fast R-CNN!
Fast R-CNN: Everything you need to know from the paper
Vision Transformer (ViT) Paper Explained
Swin Transformer V2 - Paper explained
FastV: An Image is Worth 1/2 Tokens After Layer 2
Prompt-to-Prompt (P2P) image Editing - Method Explained
GLIGEN (CVPR2023): Open-Set Grounded Text-to-Image Generation
Diffusion Models (DDPM & DDIM) - Easily explained!
MetaFormer is Actually What You Need for Vision
Swin Transformer - Paper Explained
PoseGPT (ChatPose): Chatting about 3D Human Pose
Masked Autoencoders (MAE) Paper Explained
The Entropy Enigma: Success and Failure of Entropy Minimization
Autoregressive Image Generation without Vector Quantization
Tent: Fully Test-time Adaptation by Entropy Minimization
HD-GCN (ICCV2023): Skeleton-Based Action Recognition
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
TokenHMR (CVPR2024): Advancing Human Mesh Recovery witha Tokenized Pose Representation
VPD (ICCV2023): Unleashing Text-to-Image Diffusion Models for Visual Perception
SHViT (CVPR2024): Single-Head Vision Transformer with Memory Efficient Macro Design
ConvNet beats Vision Transformers (ConvNeXt) Paper explained