GLIGEN (CVPR2023): Open-Set Grounded Text-to-Image Generation

Опубликовано: 27 Июнь 2026
на канале: Soroush Mehraban
920
23

In this video, I review the GLIGEN paper from CVPR2023 that proposes a new type of diffusion model that can receive grounding condition (bounding box, depth map, etc.) in addition to text to generate images with more constraints.

Project page: https://gligen.github.io/

Table of Content:
00:00 Intro
01:00 Grounding Tokenization
03:55 Architecture
07:21 Scheduled Sampling
09:21 Results

Icon made by Freepik from flaticon.com