OpenAI DALL·E 2: Hierarchical text conditional image generation with clip latents

Опубликовано: 04 Май 2026
на канале: Data Science Gems
1,280
14

DALL·E 2 is a 3.5B text-to-image generation model which combines CLIP, prior and diffusion decoder
It enerates diverse set of images. It generates 4x better resolution images than DALL·E and preferred by human judges more than 70% of the time both in caption matching and photorealism. DALL·E 2 works better with "longer and more detailed" input sentences.

In this video, I will briefly provide an overview of how DALL·E 2 can be used. I will also talk about training aspects of DALL·E 2. Lastly, I will compare DALL·E 2 with GLIDE and DALL·E 1, and comment on its drawbacks.

Here is the agenda:

00:00:00 What is DALL·E 2 useful for?
00:05:30 What is DALL·E 2 architecture?
00:13:12 Generating text diffs using DALL·E 2
00:16:01 DALL·E 2 Evaluation Comparison with GLIDE and DALL·E 1
00:19:14 DALL·E 2 Social and Technical Limitations

For more details, please look at https://cdn.openai.com/papers/dall-e-... and https://labs.openai.com/

For better understanding, also look at these videos:
Diffusion model based training of GLIDE:    • OpenAI GLIDE: Towards Photorealistic Image...  
DALL·E 1:    • DALL-E: Zero Shot Text to Image Generation  
CLIP:    • OpenAI CLIP: Connecting Text and Images  

Citation
Ramesh, Aditya, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. "Hierarchical text-conditional image generation with clip latents." arXiv preprint arXiv:2204.06125 (2022).