2021

MURAL: Multimodal, Multitask Retrieval Across Languages

Jain, Aashi, Guo, Mandy, Srinivasan, Krishna et al.

Understand

Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages.

  • We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Languages), a dual encoder that solves two tasks: 1) image-text matching and 2) translation pair matching.
  • By incorporating billions of translation pairs, MURAL extends ALIGN (Jia et al.
  • PMLR'21)--a state-of-the-art dual encoder learned from 1.8 billion noisy image-text pairs.

Reading the bibliography…