2020

Probing Contextual Language Models for Common Ground with Visual Representations

Ilharco, Gabriel, Zellers, Rowan, Farhadi, Ali et al.

Understand

The success of large-scale contextual language models has attracted great interest in probing what is encoded in their representations.

  • In this work, we consider a new question: to what extent contextual representations of concrete nouns are aligned with corresponding visual representations? We design a probing model that evaluates how effective are text-only representations in distinguishing between matching and non-matching visual representations.
  • Our findings show that language representations alone provide a strong signal for retrieving image patches from the correct object categories.
  • Moreover, they are effective in retrieving specific instances of image patches; textual context plays an important role in this process.

Reading the bibliography…