2020

Self-Supervised MultiModal Versatile Networks

Alayrac, Jean-Baptiste, Recasens, Adrià, Schneider, Rosalia et al.

Understand

Videos are a rich source of multi-modal supervision.

  • In this work, we learn representations using self-supervision by leveraging three modalities naturally present in videos: visual, audio and language streams.
  • To this end, we introduce the notion of a multimodal versatile network -- a network that can ingest multiple modalities and whose representations enable downstream tasks in multiple modalities.
  • In particular, we explore how best to combine the modalities, such that fine-grained representations of the visual and audio modalities can be maintained, whilst also integrating text into a common embedding.

Reading the bibliography…