Imagen: Photorealistic Text to Image Diffusion Models with Deep Language Understanding

Опубликовано: 05 Июль 2026
на канале: Data Science Gems
588
12

Imagen is a text-to-image diffusion model. It provdes better photorealism and text-image alignment compared to DALLE-2. It is a 3B parameter model with diffusion, and 2 upsamplers. It has been evaluated on COCO and Drawbench. Drawbench is a set of prompts for text-to-image model evaluation. It has been extended to Imagen Editor for image inpainting and Imagen Videos for video generation.

In this video, I will briefly provide an overview of Imagen and Drawbench.

Here is the agenda:

00:00:00 What is Imagen?
00:01:26 What is the architecture of Imagen?
00:09:40 How does Imagen perform?
00:14:55 What is DrawBench?
00:16:23 What is Imagen Editor and Imagen Video?

For more details, please look at http://imagen.research.google and https://arxiv.org/pdf/2205.11487.pdf

Saharia, Chitwan, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour et al. "Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding." In NIPS 2022