Fetching the paper…
Reading the bibliography…
Contrastive Masked Autoencoder (CMAE), as a new self-supervised framework, has shown its potential of learning expressive feature representations in visual image recognition.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Earlier work this paper cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
J. Lin, C. Gan, and S. Han · 2019
Earlier work this paper cites.
Video classification with channel-separated convolutional networks
D. Tran, H. Wang, L. Torresani, and M. Feiszli · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al · 2020
Earlier work this paper cites.
Augment your batch: Improving generalization through instance repetition
E. Hoffer, T. Ben-Nun, I. Hubara, N. Giladi, T. Hoefler, and D. Soudry · 2020
Earlier work this paper cites.
Vivit: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid · 2021
Cited alongside, same era.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani · 2021
Cited alongside, same era.
Multiscale vision transformers
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer · 2021
Cited alongside, same era.
Keeping your eye on the ball: Trajectory attention in video transformers
M. Patrick, D. Campbell, Y. Asano, I. Misra, F. Metze, C. Feichtenhofer, A. Vedaldi, and J. F. Henriques · 2021
Cited alongside, same era.
Vimpac: Video pre-training via masked token prediction and contrastive learning
Convmae: Masked convolution meets masked autoencoders
P. Gao, T. Ma, H. Li, J. Dai, and Y. Qiao · 2022
Later among the works it cites.
Omnimae: Single model masked pretraining on images and videos
R. Girdhar, A. El-Nouby, M. Singh, K. V. Alwala, A. Joulin, and I. Misra · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Contrastive masked autoencoders are stronger vision learners
Z. Huang, X. Jin, C. Lu, Q. Hou, M.-M. Cheng, D. Fu, X. Shen, and J. Feng · 2022
Later among the works it cites.
Architecture-agnostic masked image modeling–from vit back to cnn
S. Li, D. Wu, F. Wu, Z. Zang, B. Sun, H. Li, X. Xie, S. Li, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Tan, J. Lei, T. Wolf, and M. Bansal · 2021
Cited alongside, same era.
Tdn: Temporal difference networks for efficient action recognition
L. Wang, Z. Tong, B. Ji, and G. Wu · 2021
Cited alongside, same era.
Masked autoencoders as spatiotemporal learners
C. Feichtenhofer, H. Fan, Y. Li, and K. He · 2022
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He
Cited in the paper.
The" something something" video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al
Cited in the paper.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu
Cited in the paper.
Tam: Temporal adaptive module for video recognition
Z. Liu, L. Wang, W. Wu, C. Qian, and T. Lu
Cited in the paper.
Later among the works it cites.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Z. Tong, Y. Song, J. Wang, and L. Wang · 2022
Later among the works it cites.
Bevt: Bert pretraining of video transformers
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan · 2022
Later among the works it cites.