Text-to-image models use AI and a training set, often a very large one, to understand text and transform it into an image.
Imagen is one of those models, and it’s a diffusion-based approach. It generates images with a remarkable degree of photorealism and language understanding.
In this episode of Meet A Google Researcher, Sara Mahdavi chats with Chitwan Saharia and William Chan - two of the ground breaking researchers who have developed Imagen. We’ll chat about this generative model, what makes Imagen different from other text-to-image models, and where the future of Imagen lies.
Resources:
Check out Imagen → https://goo.gle/3yvFNsS
Check out the “How AI creates photorealistic images from text” blog post → https://goo.gle/3CJQ5YW
Follow William on Twitter → https://goo.gle/3ENcZQL
Follow Chitwan on Twitter → https://goo.gle/3s4q9RN
Follow Sara on Twitter → https://goo.gle/3CBr15c
Chapters:
0:00 - Intro
1:08 - Researcher Intros
1:51 - What is Imagen?
3:14 - How is Imagen different from other text-to-image models?
4:31 - Areas where Imagen succeeds and fails?
6:37 - What was your reaction to Imagen's first images?
8:03 - What did you try to get the model to work?
9:59 - Did Imagen understand its prompts right away? Did it all work at once?
11:11 - Initial pictures of Imagen.
11:54 - How did people outside of Google react to Imagen?
12:35 - The artistic process with Imagen
14:31 - How can Imagen assist artists?
15:47 - How can Imagen be improved?
17:27 - What does it take to be a Google Researcher and develop these models?
18:58 - Closing remarks
Watch more:
Watch more episodes of Meet A Google Researcher → https://goo.gle/MeetAGoogleResearcher
Subscribe to the Google Research Channel → https://goo.gle/GoogleResearch
#MeetAGoogleResearcher