Fetching the paper…

CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations · Around