TIMESTAMPS:
In this Pytorch Tutorial video we combine a vision transformer Encoder with a text Decoder to create a Model that can generate text conditioned on an Image. We show that we can use this network to generate a description of an image.
Donations, Help Support this work!
https://www.buymeacoffee.com/lukeditria
The corresponding code is available here! ( Section 14)
https://github.com/LukeDitria/pytorch...
Discord Server:
/ discord