2022

Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

Tumanyan, Narek, Geyer, Michal, Bagon, Shai et al.

Understand

Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, allowing us to synthesize diverse images that convey highly complex visual concepts.

  • However, a pivotal challenge in leveraging such models for real-world content creation tasks is providing users with control over the generated content.
  • In this paper, we present a new framework that takes text-to-image synthesis to the realm of image-to-image translation -- given a guidance image and a target text prompt, our method harnesses the power of a pre-trained text-to-image diffusion model to generate a new image that complies with the target text, while preserving the semantic layout of the source image.
  • Specifically, we observe and empirically demonstrate that fine-grained control over the generated structure can be achieved by manipulating spatial features and their self-attention inside the model.

Reading the bibliography…