2021

Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Bain, Max, Nagrani, Arsha, Varol, Gül et al.

Understand

Our objective in this work is video-text retrieval - in particular a joint embedding that enables efficient text-to-video retrieval.

  • The challenges in this area include the design of the visual architecture and the nature of the training data, in that the available large scale video-text training datasets, such as HowTo100M, are noisy and hence competitive performance is achieved only at scale through large amounts of compute.
  • We address both these challenges in this paper.
  • We propose an end-to-end trainable model that is designed to take advantage of both large-scale image and video captioning datasets.

Reading the bibliography…