CoAtNet: Marrying Convolution and Attention for All Data Sizes - Paper Explained

Опубликовано: 03 Июнь 2026
на канале: Halfling Wizard
4,156
121

In this video, I’ll try to present google brain’s paper “CoAtNet: Marrying Convolution and Attention for All Data Sizes”. This paper is another attempt to combine attention and convolution in the hopes of creating an architecture that incorporates the best features of each.

📑 Chapters:
00:00 Introduction
02:18 Merging Convolution and Self-Attention
03:45 Relative Position Representations
06:04 Vertical Layout Design
08:14 Experiments

📝 Link to the paper:
https://arxiv.org/abs/2106.04803

👥 Authors:
Zihang Dai, Hanxiao Liu, Quoc V. Le, and Mingxing Tan
🔗 Helpful Links:
My Video on the Paper "An Image Is Worth 16x16 Words"
   • An Image Is Worth 16x16 Words - Paper Expl...  
My Video on the Paper "Attention is All you Need"
   • Attention Is All You Need - Paper Explained  

🙋‍♂️ Find me on:
Find me on: halflingwizard.me

🎁 Support the Channel:
If you’d like to support my work, you can check out my wishlist here: https://www.amazon.com/registries/gl/...
Your support helps me keep creating content like this. Thank you for being part of this journey!

#Attention #transformer #vision_transformer #computer_vision #deep_learning