2020

Towards Zero-shot Cross-lingual Image Retrieval

Aggarwal, Pranav, Kale, Ajinkya

Understand

There has been a recent spike in interest in multi-modal Language and Vision problems.

  • On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual.
  • We try to bridge this gap with a zero-shot approach for learning multi-modal representations using cross-lingual pre-training on the text side.
  • We present a simple yet practical approach for building a cross-lingual image retrieval model which trains on a monolingual training dataset but can be used in a zero-shot cross-lingual fashion during inference.

Reading the bibliography…