Fetching the paper…

ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders · Around