DALL·E 2 is a 3.5B text-to-image generation model which combines CLIP, prior and diffusion decoder
It enerates diverse set of images. It generates 4x better resolution images than DALL·E and preferred by human judges more than 70% of the time both in caption matching and photorealism. DALL·E 2 works better with "longer and more detailed" input sentences.
In this video, I will briefly provide an overview of how DALL·E 2 can be used. I will also talk about training aspects of DALL·E 2. Lastly, I will compare DALL·E 2 with GLIDE and DALL·E 1, and comment on its drawbacks.
Here is the agenda:
00:00:00 What is DALL·E 2 useful for?
00:05:30 What is DALL·E 2 architecture?
00:13:12 Generating text diffs using DALL·E 2
00:16:01 DALL·E 2 Evaluation Comparison with GLIDE and DALL·E 1
00:19:14 DALL·E 2 Social and Technical Limitations
For more details, please look at https://cdn.openai.com/papers/dall-e-... and https://labs.openai.com/
For better understanding, also look at these videos:
Diffusion model based training of GLIDE: • OpenAI GLIDE: Towards Photorealistic Image...
DALL·E 1: • DALL-E: Zero Shot Text to Image Generation
CLIP: • OpenAI CLIP: Connecting Text and Images
Citation
Ramesh, Aditya, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. "Hierarchical text-conditional image generation with clip latents." arXiv preprint arXiv:2204.06125 (2022).