Stable Diffusion Explained - How text to image generation works using U-Net Noise Predictor

Опубликовано: 28 Октябрь 2024
на канале: Architecture Bytes
1,048
25

#stablediffusion #aiimagegenerator #artificialintelligence #neuralnetworks #midjourneyai #dalle3 #chatgpt
Let's learn how Latent Stable Diffusion works to generate images based on text prompts. It uses U-Net as the Noise Predictor. We will see the Architecture and Steps involved in the process.
Stable Diffusion Architecture

Timings:
00:10 High Level Design
00:33 Diffusion
1:08 Training & Inference
2:26 Latent Space - VAE
3:36 Architecture
4:28 CLIP Text Encoder
6:50 U-Net