2020

VirTex: Learning Visual Representations from Textual Annotations

Desai, Karan, Johnson, Justin

Understand

The de-facto approach to many vision tasks is to start from pretrained visual representations, typically learned via supervised training on ImageNet.

  • Recent methods have explored unsupervised pretraining to scale to vast quantities of unlabeled images.
  • In contrast, we aim to learn high-quality visual representations from fewer images.
  • To this end, we revisit supervised pretraining, and seek data-efficient alternatives to classification-based pretraining.

Reading the bibliography…