2021

LAFITE: Towards Language-Free Training for Text-to-Image Generation

Zhou, Yufan, Zhang, Ruiyi, Chen, Changyou et al.

Understand

One of the major challenges in training text-to-image generation models is the need of a large number of high-quality image-text pairs.

  • While image samples are often easily accessible, the associated text descriptions typically require careful human captioning, which is particularly time- and cost-consuming.
  • In this paper, we propose the first work to train text-to-image generation models without any text data.
  • Our method leverages the well-aligned multi-modal semantic space of the powerful pre-trained CLIP model: the requirement of text-conditioning is seamlessly alleviated via generating text features from image features.

Reading the bibliography…