2019

Self-Supervised Learning by Cross-Modal Audio-Video Clustering

Alwassel, Humam, Mahajan, Dhruv, Korbar, Bruno et al.

Understand

Visual and audio modalities are highly correlated, yet they contain different information.

  • Their strong correlation makes it possible to predict the semantics of one from the other with good accuracy.
  • Their intrinsic differences make cross-modal prediction a potentially more rewarding pretext task for self-supervised learning of video and audio representations compared to within-modality learning.
  • Based on this intuition, we propose Cross-Modal Deep Clustering (XDC), a novel self-supervised method that leverages unsupervised clustering in one modality (e.g., audio) as a supervisory signal for the other modality (e.g., video).

Reading the bibliography…