Fetching the paper…
Reading the bibliography…
Self-supervised pre-training for images without labels has recently achieved promising performance in image classification.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Erhan, D., Courville, A., Bengio, Y., and Vincent, P · 2010
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Efficient inference in occlusion-aware generative models of images
Huang, J. and Murphy, K · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, À., Oliva, A., and Torralba, A · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Lookahead optimizer: k steps forward, 1 step back
Zhang, M. R., Lucas, J., Ba, J., and Hinton, G. E · 2019
Cited alongside, same era.
Quantifying attention flow in transformers
Abnar, S. and Zuidema, W. H · 2020
Cited alongside, same era.
Bootstrap your own latent - a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., Piot, B., kavukcuoglu, k., Munos, R., and Valko, M · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. B · 2020
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Are large-scale datasets necessary for self-supervised pre-training?
El-Nouby, A., Izacard, G., Touvron, H., Laptev, I., Jégou, H., and Grave, E · 2021
Later among the works it cites.
Xcit: Cross-covariance image transformers
El-Nouby, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., and Jégou, H · 2021
Later among the works it cites.
GENESIS-V2: inferring unordered object representations without iterative refinement
Engelcke, M., Jones, O. P., and Posner, I · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Cited alongside, same era.
BEiT: BERT pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Cited alongside, same era.
Peco: Perceptual codebook for BERT pre-training of vision transformers
Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., and Yu, N · 2021
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G
Cited in the paper.
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. B · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Masked feature prediction for self-supervised visual pre-training
Wei, C., Fan, H., Xie, S., Wu, C., Yuille, A. L., and Feichtenhofer, C · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Tay, F. E., Feng, J., and Yan, S · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.