2022

DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models

Cho, Jaemin, Zala, Abhay, Bansal, Mohit

Understand

Recently, DALL-E, a multimodal transformer language model, and its variants, including diffusion models, have shown high-quality text-to-image generation capabilities.

  • However, despite the realistic image generation results, there has not been a detailed analysis of how to evaluate such models.
  • In this work, we investigate the visual reasoning capabilities and social biases of different text-to-image models, covering both multimodal transformer language models and diffusion models.
  • First, we measure three visual reasoning skills: object recognition, object counting, and spatial relation understanding.

Reading the bibliography…