Stability AI has detailed Stable Cascade, a new image generation model built on the Würstchen architecture, which improves performance and accuracy compared to the SDXL architecture previously used. Quoting Venturebeat:
Stability AI has been steadily iterating on its core Stable Diffusion model since 2022. The SDXL 1.0 release in July 2023 marked a new flagship release, which was further accelerated with the SDXL Turbo update in November 2023.
Stable Cascade uses somewhat of a different architecture than SDXL to generate images that Stability AI researchers hope will be more efficient. The new approach builds on the Würstchen architecture, which uses a series of innovative techniques to improve performance and accuracy.
“A key contribution of our work is to develop a latent diffusion technique in which we learn a detailed but extremely compact semantic image representation used to guide the diffusion process,” the Würstchen research abstract states. “This highly compressed representation of an image provides much more detailed guidance compared to latent representations of language and this significantly reduces the computational requirements to achieve state-of-the-art results.”
Unlike Stable Diffusion which uses a single large model, Stable Cascade utilizes a pipeline of three distinct smaller models referred to as Stages A, B and C. This modular architecture provides major advantages in training efficiency and customization.
The first stage, Stage C, transforms text prompts into compact 24×24 pixel latents. Stages A and B then decode these latents into full high-resolution images. By separating the text-to-image generation from the image decoding, the initial text-conditional model can be trained and fine-tuned much more efficiently. According to Stability AI, fine-tuning Stage C alone provides a 16x cost reduction compared to fine-tuning an equivalently sized single Stable Diffusion model.
There is also the potential for Direct Preference Optimization (DPO) to further improve image quality. In a 2023 interview with VentureBeat, Stability AI founder and CEO Emad Mostaque explained that DPO is an alternative approach to reinforcement learning used in models to tune them to human preferences. ENDQUOTE
#ainews #stabilityai #technews