2024

MaxFusion: Plug&Play Multi-Modal Generation in Text-to-Image Diffusion Models

Nair, Nithin Gopalakrishnan, Valanarasu, Jeya Maria Jose, Patel, Vishal M

Understand

Large diffusion-based Text-to-Image (T2I) models have shown impressive generative powers for text-to-image generation as well as spatially conditioned image generation.

  • For most applications, we can train the model end-toend with paired data to obtain photorealistic generation quality.
  • However, to add an additional task, one often needs to retrain the model from scratch using paired data across all modalities to retain good generation performance.
  • In this paper, we tackle this issue and propose a novel strategy to scale a generative model across new tasks with minimal compute.

Reading the bibliography…