Fetching the paper…

Video Representation Learning with Joint-Embedding Predictive Architectures · Around