2022

Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Feng, Weixi, He, Xuehai, Fu, Tsu-Jui et al.

Understand

Large-scale diffusion models have achieved state-of-the-art results on text-to-image synthesis (T2I) tasks.

  • Despite their ability to generate high-quality yet creative images, we observe that attribution-binding and compositional capabilities are still considered major challenging issues, especially when involving multiple objects.
  • In this work, we improve the compositional skills of T2I models, specifically more accurate attribute binding and better image compositions.
  • To do this, we incorporate linguistic structures with the diffusion guidance process based on the controllable properties of manipulating cross-attention layers in diffusion-based T2I models.

Reading the bibliography…