2022

Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?

Yun, Tian, Bhalla, Usha, Pavlick, Ellie et al.

Understand

Vision-language (VL) pretrained models have achieved impressive performance on multimodal reasoning and zero-shot recognition tasks.

  • Many of these VL models are pretrained on unlabeled image and caption pairs from the internet.
  • In this paper, we study whether representations of primitive concepts--such as colors, shapes, or the attributes of object parts--emerge automatically within these pretrained VL models.
  • We propose a two-step framework, Compositional Concept Mapping (CompMap), to investigate this.

Reading the bibliography…