Captioning Images with a Transformer, from Scratch! PyTorch Deep Learning Tutorial

Опубликовано: 23 Май 2026
на канале: Luke Ditria
6,413
158

TIMESTAMPS:

In this Pytorch Tutorial video we combine a vision transformer Encoder with a text Decoder to create a Model that can generate text conditioned on an Image. We show that we can use this network to generate a description of an image.

Donations, Help Support this work!
https://www.buymeacoffee.com/lukeditria

The corresponding code is available here! ( Section 14)
https://github.com/LukeDitria/pytorch...

Discord Server:
  / discord