Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, allowing us to synthesize diverse images that convey highly complex visual concepts. However, a pivotal challenge in leveraging such models for real-world content creation tasks is providing users with control over the generated content. In this paper, we present a new framework that takes text-to-image synthesis to the realm of image-to-image translation — given a guidance image and a target text prompt, our method harnesses the power of a pre-trained text-to-image diffusion model to generate a new image that complies with the target text, while preserving the semantic layout of the source image. Specifically, we observe and empirically demonstrate that fine-grained control over the generated structure can be achieved by manipulating spatial features and their self-attention inside the model.
Michal Geyer and Narek Tumanya are Masters students at the Weizmann Institute of Science in the Computer Vision department.
Join the Computer Vision Meetup friendliest to your timezone:
https://www.meetup.com/pro/computer-v...
Recorded on June 8, 2023 at the virtual Computer Vision Meetup.
#computervision #machinelearning #datascience #ai