2023

P+: Extended Textual Conditioning in Text-to-Image Generation

Voynov, Andrey, Chu, Qinghao, Cohen-Or, Daniel et al.

Understand

We introduce an Extended Textual Conditioning space in text-to-image models, referred to as $P+$.

  • This space consists of multiple textual conditions, derived from per-layer prompts, each corresponding to a layer of the denoising U-net of the diffusion model.
  • We show that the extended space provides greater disentangling and control over image synthesis.
  • We further introduce Extended Textual Inversion (XTI), where the images are inverted into $P+$, and represented by per-layer tokens.

Reading the bibliography…