In this video, I'll be discussing the paper "Show, Attend, and Tell: Neural Image Caption Generation with Visual Attention."
This paper created captions using the attention method, and it was considered state-of-the-art at the time of publication. It introduces a modular encoder-decoder with a CNN as the encoder and an RNN as the decoder, as well as an attention mechanism; along with A new, “hard” approach to calculating the attention weight has been proposed, which essentially lowers the computational cost.
📑 Chapters:
0:00 Abstract
0:56 Introduction
2:19 Model Details
6:16 "Hard" and "Soft" Attention
8:41 Results
10:42 Conclusion
📝 Link to the paper:
https://arxiv.org/abs/1502.03044
👥 Authors:
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio
🔗 Helpful Links:
Toward data science’s article ‘Illustrated Guide to LSTM’s and GRU’s: A step by step explanation’:
https://towardsdatascience.com/illust...
🙏 I'd like to express my gratitude to Dr. Nasersharif, my supervisor, for suggesting this paper to me.
🙋♂️ Find me on: halflingwizard.me
🎁 Support the Channel:
If you’d like to support my work, you can check out my wishlist here: https://www.amazon.com/registries/gl/...
Your support helps me keep creating content like this. Thank you for being part of this journey!
#Attention #Image_captioning #computer_vision #deep_learning