Fetching the paper…
Reading the bibliography…
Masked autoencoders are scalable vision learners, as the title of MAE \cite{he2022masked}, which suggests that self-supervised learning (SSL) in vision might undertake a similar trajectory as in NLP.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in CVPR , 2005
2005
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in ICML , 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, P.-A. Manzagol, and L. Bottou, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion.” Journal of machine learning research , 2010
2010
Earlier work this paper cites.
A. Ng et al. , “Sparse autoencoder,” CS294A Lecture notes , 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in NeurIPS , 2012
2012
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in CVPR , 2016
2016
Earlier work this paper cites.
M. Noroozi and P. Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” in ECCV , 2016
2016
Earlier work this paper cites.
G. Larsson, M. Maire, and G. Shakhnarovich, “Learning representations for automatic colorization,” in ECCV , 2016
2016
Earlier work this paper cites.
R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” in ECCV , 2016
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
——, “Colorization as a proxy task for visual understanding,” in CVPR , 2017
2017
Earlier work this paper cites.
——, “Split-brain autoencoders: Unsupervised learning by cross-channel prediction,” in CVPR , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in CVPR , 2017
2017
Earlier work this paper cites.
W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” in NeurIPS , 2017
2017
Earlier work this paper cites.
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in ICLR , 2017
2017
Earlier work this paper cites.
S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” ICLR , 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in CVPR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio, “Graph attention networks,” ICLR , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” NAACL , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , 2019
2019
Earlier work this paper cites.
A. Bilogur, “Notes on gpt-2 and bert models,” Kaggle blog , 2019
2019
Earlier work this paper cites.
K. Xu, W. Hu, J. Leskovec, and S. Jegelka, “How powerful are graph neural networks?” ICLR , 2019
2019
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML , 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , 2020
2020
Earlier work this paper cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in ICML , 2020
2020
Earlier work this paper cites.
——, “Generative pretraining from pixels,” in OpenAI blog , 2020
2020
Earlier work this paper cites.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Advances in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
I. Misra and L. v. d. Maaten, “Self-supervised learning of pretext-invariant representations,” in CVPR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas, “Reinforcement learning with augmented data,” Advances in neural information processing systems , vol. 33, pp. 19 884–19 895, 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in ICML , 2021
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny, “Barlow twins: Self-supervised learning via redundancy reduction,” ICML , 2021
2021
Cited alongside, same era.
X. Chen and K. He, “Exploring simple siamese representation learning,” in CVPR , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” 2021
2021
Cited alongside, same era.
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick, “Early convolutions help transformers see better,” NeurIPS , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
X. Chen, S. Xie, and K. He, “An empirical study of training self-supervised vision transformers,” ICCV , 2021
2021
Cited alongside, same era.
Z. Li, Z. Chen, F. Yang, W. Li, Y. Zhu, C. Zhao, R. Deng, L. Wu, R. Zhao, M. Tang et al. , “Mst: Masked self-supervised transformer for visual representation,” NeurIPS , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
L. Jing, J. Zhu, and Y. LeCun, “Masked siamese convnets,” arXiv preprint arXiv:2206.07700 , 2022
2022
Closest in time.
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong, “ibot: Image bert pre-training with online tokenizer,” ICLR , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
J. An, Y. Bai, H. Chen, Z. Gao, and G. Litjens, “Masked autoencoders pre-training in multiple instance learning for whole slide image classification,” in Medical Imaging with Deep Learning , 2022
2022
Closest in time.
2022
Closest in time.
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan, “Bevt: Bert pretraining of video transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 14 733–14 743
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Z.-Y. Dou, Y. Xu, Z. Gan, J. Wang, S. Wang, L. Wang, C. Zhu, P. Zhang, L. Yuan, N. Peng et al. , “An empirical study of training end-to-end vision-and-language transformers,” in CVPR , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
C. Min, D. Zhao, L. Xiao, Y. Nie, and B. Dai, “Voxel-mae: Masked autoencoders for pre-training large-scale point clouds,” arXiv e-prints , pp. arXiv–2206, 2022
2022
Closest in time.
2022
Closest in time.
H. Chen, S. Zhang, and G. Xu, “Graph masked autoencoder,” arXiv preprint arXiv:2202.08391 , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
M. Zha, “Time series generation with masked autoencoder,” arXiv preprint arXiv:2201.07006 , 2022
2022
Closest in time.
2022
Closest in time.