2021

SLIP: Self-supervision meets Language-Image Pre-training

Mu, Norman, Kirillov, Alexander, Wagner, David et al.

Understand

Recent work has shown that self-supervised pre-training leads to improvements over supervised learning on challenging visual recognition tasks.

  • CLIP, an exciting new approach to learning with language supervision, demonstrates promising performance on a wide variety of benchmarks.
  • In this work, we explore whether self-supervised learning can aid in the use of language supervision for visual representation learning.
  • We introduce SLIP, a multi-task learning framework for combining self-supervised learning and CLIP pre-training.

Reading the bibliography…