Fetching the paper…
Reading the bibliography…
Recently, self-supervised Masked Autoencoders (MAE) have attracted unprecedented attention for their impressive representation learning ability.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D.; and Dietterich, T. 2019 · 1903
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
The pascal visual object classes (VOC) challenge
Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010 · 2010
Earlier work this paper cites.
The German traffic sign recognition benchmark: a multi-class classification competition
Stallkamp, J.; Schlipsing, M.; Salmen, J.; and Igel, C. 2011 · 2011
Earlier work this paper cites.
The MNIST database of handwritten digit images for machine learning research
Deng, L. 2012 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
A new performance measure and evaluation benchmark for road detection algorithms
Fritsch, J.; Kuehnl, T.; and Geiger, A. 2013 · 2013
Earlier work this paper cites.
FER 2013: Kaggle challenges in representation learning facial expression recognition
kaggle. 2013 · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Gool, L. V. 2014 · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M.; Maji, S.; Kokkinos, I.; Mohamed, S.; and Vedaldi, A. 2014 · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R.; Donahue, J.; Darrell, T.; and Malik, J. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015 · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Long, J.; Shelhamer, E.; and Darrell, T. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Augmenting supervised neural networks with unsupervised objectives for large-scale image classification
Zhang, Y.; Lee, K.; and Lee, H. 2016 · 2016
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
Supervised contrastive learning
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020 · 2020
Later among the works it cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D.; Firooz, H.; Mohan, A.; Goswami, V.; Singh, A.; Ringshia, P.; and Testuggine, D. 2020 · 2020
Later among the works it cites.
MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark
MMSegmentation, C. 2020 · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
Bao, H.; Dong, L.; and Wei, F. 2021 · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheng, G.; Han, J.; and Lu, X. 2017 · 2017
Cited alongside, same era.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Normface: L2 hypersphere embedding for face verification
Wang, F.; Xiang, X.; Cheng, J.; and Yuille, A. L. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Supervised autoencoders: Improving generalization performance with unsupervised regularizers
Le, L.; Patterson, A.; and White, M. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Cited alongside, same era.
Chen*, X.; Xie*, S.; and He, K. 2021 · 2021
Later among the works it cites.
Peco: Perceptual codebook for bert pre-training of vision transformers
Dong, X.; Bao, J.; Zhang, T.; Chen, D.; Zhang, W.; Yuan, L.; Chen, D.; Wen, F.; and Yu, N. 2021 · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A.; Kolesnikov, A.; Zhai, X.; Wightman, R.; Uszkoreit, J.; and Beyer, L. 2021 · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Later among the works it cites.
Masked Feature Prediction for Self-Supervised Visual Pre-Training
Wei, C.; Fan, H.; Xie, S.; Wu, C.-Y.; Yuille, A.; and Feichtenhofer, C. 2021 · 2021
Later among the works it cites.
Masked Siamese Networks for Label-Efficient Learning
Assran, M.; Caron, M.; Misra, I.; Bojanowski, P.; Bordes, F.; Vincent, P.; Joulin, A.; Rabbat, M.; and Ballas, N. 2022 · 2022
Closest in time.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A.; Hsu, W.-N.; Xu, Q.; Babu, A.; Gu, J.; and Auli, M. 2022 · 2022
Closest in time.
Contrastive Masked Autoencoders are Stronger Vision Learners
Huang, Z.; Jin, X.; Lu, C.; Hou, Q.; Cheng, M.-M.; Fu, D.; Shen, X.; and Feng, J. 2022 · 2022
Closest in time.
Touvron, H.; Cord, M.; and Jégou, H. 2022 · 2022
Closest in time.
Repre: Improving self-supervised vision transformer with reconstructive pre-training
Wang, L.; Liang, F.; Li, Y.; Ouyang, W.; Zhang, H.; and Shao, J. 2022 · 2022
Closest in time.