In this video, I review the GLIGEN paper from CVPR2023 that proposes a new type of diffusion model that can receive grounding condition (bounding box, depth map, etc.) in addition to text to generate images with more constraints.
Project page: https://gligen.github.io/
Table of Content:
00:00 Intro
01:00 Grounding Tokenization
03:55 Architecture
07:21 Scheduled Sampling
09:21 Results
Icon made by Freepik from flaticon.com