2020

Video Understanding as Machine Translation

Korbar, Bruno, Petroni, Fabio, Girdhar, Rohit et al.

Understand

With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of video representations.

  • Most prior work formulates the objective as a contrastive metric learning problem between the modalities.
  • To enable effective learning, however, these strategies require a careful selection of positive and negative samples often combined with hand-designed curriculum policies.
  • In this work we remove the need for negative sampling by taking a generative modeling approach that poses the objective as a translation problem between modalities.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…