Fetching the paper…

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training · Around