2017

Image Pivoting for Learning Multilingual Multimodal Representations

Gella, Spandana, Sennrich, Rico, Keller, Frank et al.

Understand

In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding.

  • Our model learns a common representation for images and their descriptions in two different languages (which need not be parallel) by considering the image as a pivot between two languages.
  • We introduce a new pairwise ranking loss function which can handle both symmetric and asymmetric similarity between the two modalities.
  • We evaluate our models on image-description ranking for German and English, and on semantic textual similarity of image descriptions in English.

Reading the bibliography…