Understand
With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of video representations.
- Most prior work formulates the objective as a contrastive metric learning problem between the modalities.
- To enable effective learning, however, these strategies require a careful selection of positive and negative samples often combined with hand-designed curriculum policies.
- In this work we remove the need for negative sampling by taking a generative modeling approach that poses the objective as a translation problem between modalities.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…