2023

What You See is What You Read? Improving Text-Image Alignment Evaluation

Yarom, Michal, Bitton, Yonatan, Changpinyo, Soravit et al.

Understand

Automatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative text-to-image and image-to-text tasks.

  • In this work, we study methods for automatic text-image alignment evaluation.
  • We first introduce SeeTRUE: a comprehensive evaluation set, spanning multiple datasets from both text-to-image and image-to-text generation tasks, with human judgements for whether a given text-image pair is semantically aligned.
  • We then describe two automatic methods to determine alignment: the first involving a pipeline based on question generation and visual question answering models, and the second employing an end-to-end classification approach by finetuning multimodal pretrained models.

Reading the bibliography…