Fetching the paper…
Reading the bibliography…
The use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization, in: Proceedings of the IEEE/CVF International Conference on Computer Vision
Krause, J., Stark, M., Deng, J., Fei-Fei, L., 2013 · 2013
Earlier work this paper cites.
Fine-Grained Visual Classification of Aircraft
Maji, S., Kannala, J., Rahtu, E., Blaschko, M., Vedaldi, A., 2013 · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests, in: Proceedings of the European Conference on Computer Vision
Bossard, L., Guillaumin, M., Van Gool, L., 2014 · 2014
Earlier work this paper cites.
Unsupervised Visual Representation Learning by Context Prediction, in: Proceedings of the IEEE/CVF International Conference on Computer Vision
Doersch, C., Gupta, A., Efros, A.A., 2015 · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos, in: Proceedings of the IEEE/CVF International Conference on Computer Vision
Wang, X., Gupta, A., 2015 · 2015
Earlier work this paper cites.
Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles, in: Proceedings of the European Conference on Computer Vision
Noroozi, M., Favaro, P., 2016 · 2016
Earlier work this paper cites.
Colorful image colorization, in: Proceedings of the European Conference on Computer Vision
Zhang, R., Isola, P., Efros, A.A., 2016 · 2016
Earlier work this paper cites.
Scene Parsing Through ADE20K Dataset, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A., 2017 · 2017
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations, in: Proceedings of the International Conference on Learning Representations
Gidaris, S., Singh, P., Komodakis, N., 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A.v.d., Li, Y., Vinyals, O., 2018 · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding, in: Proceedings of the European Conference on Computer Vision
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., Sun, J., 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, in: Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019 · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, in: Proceedings of the International Conference on Learning Representations
Loshchilov, I., Hutter, F., 2019 · 2019
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding, in: Advances in Neural Information Processing Systems
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R.R., Le, Q.V., 2019 · 2019
Cited alongside, same era.
Semantic Understanding of Scenes Through the ADE20K Dataset
Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., Torralba, A., 2019 · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners, in: Advances in Neural Information Processing Systems
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning, in: Advances in Neural Information Processing Systems
Grill, J.B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M., 2020 · 2020
BEiT: BERT pre-training of image Transformers, in: Proceedings of the International Conference on Learning Representations
Bao, H., Dong, L., Piao, S., Wei, F., 2022 · 2022
Later among the works it cites.
Masked Autoencoders Are Scalable Vision Learners, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R., 2022 · 2022
Later among the works it cites.
MILAN: Masked Image Pretraining on Language Assisted Representation
Hou, Z., Sun, F., Chen, Y.K., Xie, Y., Kung, S.Y., 2022 · 2022
Later among the works it cites.
Swin Transformer V2: Scaling Up Capacity and Resolution, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al., 2022 · 2022
Later among the works it cites.
BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
He, K., Fan, H., Wu, Y., Xie, S., Girshick, R., 2020 · 2020
Cited alongside, same era.
MPNet: Masked and Permuted Pre-training for Language Understanding, in: Advances in Neural Information Processing Systems
Song, K., Tan, X., Qin, T., Lu, J., Liu, T.Y., 2020 · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vision
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A., 2021 · 2021
Cited alongside, same era.
Exploring simple siamese representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Chen, X., He, K., 2021 · 2021
Cited alongside, same era.
An Empirical Study of Training Self-Supervised Vision Transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vision
Chen, X., Xie, S., He, K., 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, in: Proceedings of the International Conference on Learning Representations
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N., 2021 · 2021
Cited alongside, same era.
Are Large-scale Datasets Necessary for Self-Supervised Pre-training?
El-Nouby, A., Izacard, G., Touvron, H., Laptev, I., Jegou, H., Grave, E., 2021 · 2021
Cited alongside, same era.
Peng, Z., Dong, L., Bao, H., Ye, Q., Wei, F., 2022 · 2022
Later among the works it cites.
Beyond masking: Demystifying token-based pre-training for vision transformers
Tian, Y., Xie, L., Fang, J., Shi, M., Peng, J., Zhang, X., Jiao, J., Tian, Q., Ye, Q., 2022 · 2022
Later among the works it cites.
SimMIM: a Simple Framework for Masked Image Modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., Hu, H., 2022 · 2022
Later among the works it cites.
CAE v2: Context Autoencoder with CLIP Target
Zhang, X., Chen, J., Yuan, J., Chen, Q., Wang, J., Wang, X., Han, S., Chen, X., Pi, J., Yao, K., Han, J., Ding, E., Wang, J., 2022 · 2022
Later among the works it cites.
Context Autoencoder for Self-Supervised Representation Learning
Chen, X., Ding, M., Wang, X., Xin, Y., Mo, S., Wang, Y., Han, S., Luo, P., Zeng, G., Wang, J., 2023 · 2023
Closest in time.
PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers, in: Proceedings of the AAAI Conference on Artificial Intelligence
Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., Yu, N., 2023 · 2023
Closest in time.
Corrupted image modeling for self-supervised visual pre-training, in: Proceedings of the International Conference on Learning Representations
Fang, Y., Dong, L., Bao, H., Wang, X., Wei, F., 2023 · 2023
Closest in time.
Contrastive Masked Autoencoders are Stronger Vision Learners
Huang, Z., Jin, X., Lu, C., Hou, Q., Cheng, M.M., Fu, D., Shen, X., Feng, J., 2023 · 2023
Closest in time.
Integrally pre-trained transformer pyramid networks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Tian, Y., Xie, L., Wang, Z., Wei, L., Zhang, X., Jiao, J., Wang, Y., Tian, Q., Ye, Q., 2023 · 2023
Closest in time.
Hivit: Hierarchical vision transformer meets masked image modeling
Zhang, X., Tian, Y., Huang, W., Ye, Q., Dai, Q., Xie, L., Tian, Q., 2023 · 2023
Closest in time.