2023

Adding Conditional Control to Text-to-Image Diffusion Models

Zhang, Lvmin, Rao, Anyi, Agrawala, Maneesh

Understand

We present ControlNet, a neural network architecture to add spatial conditioning controls to large, pretrained text-to-image diffusion models.

  • ControlNet locks the production-ready large diffusion models, and reuses their deep and robust encoding layers pretrained with billions of images as a strong backbone to learn a diverse set of conditional controls.
  • The neural architecture is connected with "zero convolutions" (zero-initialized convolution layers) that progressively grow the parameters from zero and ensure that no harmful noise could affect the finetuning.
  • We test various conditioning controls, eg, edges, depth, segmentation, human pose, etc, with Stable Diffusion, using single or multiple conditions, with or without prompts.

Reading the bibliography…