Self-Attention in transfomers - Part 2

Опубликовано: 16 Август 2026
на канале: AI Bites
11,782
213

This is video no. 2 in the 5 part video series on Transformers Neural Network Architecture. This video starts with the idea of Query, Key and Value and moves on to show how it is leveraged in the Transformers architecture.

⌚️ ⌚️ ⌚️ TIMESTAMPS ⌚️ ⌚️ ⌚️
0:00 - Intro to Query, Key and Value
0:44 - Scaled Dot-Product Attention (Self-Attention)
1:46 - Learning Self-Attention
2:29 - Multi-Head Attention
3:46 - Use of Attention in Transformers
4:49 - Masking
5:27 - Advantages of Self-Attention

🛠 🛠 🛠 MY SOFTWARE TOOLS 🛠 🛠 🛠
✍️ Notion - https://affiliate.notion.so/aibites-yt
📹 OBS Studio for video editing - https://obsproject.com
📼 Manim for some animations - https://www.manim.community
🎵 My music - https://www.bensound.com and

📚 📚 📚 BOOKS I HAVE READ, REFER AND RECOMMEND 📚 📚 📚
📖 Deep Learning by Ian Goodfellow - https://amzn.to/3Wnyixv
📙 Pattern Recognition and Machine Learning by Christopher M. Bishop - https://amzn.to/3ZVnQQA
📗 Machine Learning: A Probabilistic Perspective by Kevin Murphy - https://amzn.to/3kAqThb
📘 Multiple View Geometry in Computer Vision by R Hartley and A Zisserman - https://amzn.to/3XKVOWi

MY KEY LINKS
YouTube:    / aibites​  
Twitter:   / ai_bites​  
Patreon:   / ai_bites​  
Github: https://github.com/ai-bites​

WHO AM I?
I am a Machine Learning Researcher / Practioner who has seen the grind of academia, start-ups. I started my career as a software engineer 15 years back. Because of my love for Mathematics coupled with a glimmer of luck, I graduated with a Master's in Computer Vision and Robotics in 2016 when the now happening AI revolution just started.

#transformers #machinelearning #deeplearning #aibites