2022

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Liang, Weixin, Zhang, Yuhui, Kwon, Yongchan et al.

Understand

We present modality gap, an intriguing geometric phenomenon of the representation space of multi-modal models.

  • Specifically, we show that different data modalities (e.g.
  • images and text) are embedded at arm's length in their shared representation in multi-modal models such as CLIP.
  • Our systematic analysis demonstrates that this gap is caused by a combination of model initialization and contrastive learning optimization.

Reading the bibliography…