Show, Attend And Tell - Paper Explained

Опубликовано: 20 Апрель 2026
на канале: Halfling Wizard
6,169
147

In this video, I'll be discussing the paper "Show, Attend, and Tell: Neural Image Caption Generation with Visual Attention."
This paper created captions using the attention method, and it was considered state-of-the-art at the time of publication. It introduces a modular encoder-decoder with a CNN as the encoder and an RNN as the decoder, as well as an attention mechanism; along with A new, “hard” approach to calculating the attention weight has been proposed, which essentially lowers the computational cost.

📑 Chapters:
0:00 Abstract
0:56 Introduction
2:19 Model Details
6:16 "Hard" and "Soft" Attention
8:41 Results
10:42 Conclusion

📝 Link to the paper:
https://arxiv.org/abs/1502.03044

👥 Authors:
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio
🔗 Helpful Links:
Toward data science’s article ‘Illustrated Guide to LSTM’s and GRU’s: A step by step explanation’:
https://towardsdatascience.com/illust...
🙏 I'd like to express my gratitude to Dr. Nasersharif, my supervisor, for suggesting this paper to me.

🙋‍♂️ Find me on: halflingwizard.me

🎁 Support the Channel:
If you’d like to support my work, you can check out my wishlist here: https://www.amazon.com/registries/gl/...
Your support helps me keep creating content like this. Thank you for being part of this journey!

#Attention #Image_captioning #computer_vision #deep_learning