2019

Large-scale weakly-supervised pre-training for video action recognition

Ghadiyaram, Deepti, Feiszli, Matt, Tran, Du et al.

Understand

Current fully-supervised video datasets consist of only a few hundred thousand videos and fewer than a thousand domain-specific labels.

  • This hinders the progress towards advanced video architectures.
  • This paper presents an in-depth study of using large volumes of web videos for pre-training video models for the task of action recognition.
  • Our primary empirical finding is that pre-training at a very large scale (over 65 million videos), despite on noisy social-media videos and hashtags, substantially improves the state-of-the-art on three challenging public action recognition datasets.

Reading the bibliography…