2020

Spatiotemporal Contrastive Video Representation Learning

Qian, Rui, Meng, Tianjian, Gong, Boqing et al.

Understand

We present a self-supervised Contrastive Video Representation Learning (CVRL) method to learn spatiotemporal visual representations from unlabeled videos.

  • Our representations are learned using a contrastive loss, where two augmented clips from the same short video are pulled together in the embedding space, while clips from different videos are pushed away.
  • We study what makes for good data augmentations for video self-supervised learning and find that both spatial and temporal information are crucial.
  • We carefully design data augmentations involving spatial and temporal cues.

Reading the bibliography…