2020

Memory-augmented Dense Predictive Coding for Video Representation Learning

Han, Tengda, Xie, Weidi, Zisserman, Andrew

Understand

The objective of this paper is self-supervised learning from video, in particular for representations for action recognition.

  • We make the following contributions: (i) We propose a new architecture and learning framework Memory-augmented Dense Predictive Coding (MemDPC) for the task.
  • It is trained with a predictive attention mechanism over the set of compressed memories, such that any future states can always be constructed by a convex combination of the condense representations, allowing to make multiple hypotheses efficiently.
  • (ii) We investigate visual-only self-supervised video representation learning from RGB frames, or from unsupervised optical flow, or both.

Reading the bibliography…