2021

Cross-Modal Retrieval Augmentation for Multi-Modal Classification

Gur, Shir, Neverova, Natalia, Stauffer, Chris et al.

Understand

Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing.

  • Here, we explore the use of unstructured external knowledge sources of images and their corresponding captions for improving visual question answering (VQA).
  • First, we train a novel alignment model for embedding images and captions in the same space, which achieves substantial improvement in performance on image-caption retrieval w.r.t.
  • similar methods.

Reading the bibliography…