2024

No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models

Pouget, Angéline, Beyer, Lucas, Bugliarello, Emanuele et al.

Understand

We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs).

  • Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings.
  • First, the common filtering of training data to English image-text pairs disadvantages communities of lower socioeconomic status and negatively impacts cultural understanding.
  • Notably, this performance gap is not captured by - and even at odds with - the currently popular evaluation metrics derived from the Western-centric ImageNet and COCO datasets.

Reading the bibliography…