2022

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Balaji, Yogesh, Nah, Seungjun, Huang, Xun et al.

Understand

Large-scale diffusion-based generative models have led to breakthroughs in text-conditioned high-resolution image synthesis.

  • Starting from random noise, such text-to-image diffusion models gradually synthesize images in an iterative fashion while conditioning on text prompts.
  • We find that their synthesis behavior qualitatively changes throughout this process: Early in sampling, generation strongly relies on the text prompt to generate text-aligned content, while later, the text conditioning is almost entirely ignored.
  • This suggests that sharing model parameters throughout the entire generation process may not be ideal.

Reading the bibliography…