Fetching the paper…
Reading the bibliography…
Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. et al · 2009
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
Cortes, C., Mohri, M., and Rostamizadeh, A · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V · 2012
Earlier work this paper cites.
Feature selection via dependence maximization
Song, L., Smola, A., Gretton, A., Bedo, J., and Borgwardt, K · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A · 2013
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M. and Favaro, P · 2016
Earlier work this paper cites.
Colorful image colorization
Zhang, R., Isola, P., and Efros, A. A · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Gidaris, S., Singh, P., and Komodakis, N · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Earlier work this paper cites.
The inaturalist species classification and detection dataset
Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Cited alongside, same era.
On the efficacy of knowledge distillation
Cho, J. H. and Hariharan, B · 2019
Cited alongside, same era.
Searching for mobilenetv3
Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., and Adam, H · 2019
Cited alongside, same era.
Knowledge distillation via route constrained optimization
Jin, X., Peng, B., Wu, Y., Liu, Y., Liu, J., Liang, D., Yan, J., and Hu, X · 2019
Cited alongside, same era.
Large-scale long-tailed recognition in an open world
Liu, Z., Miao, Z., Zhan, X., Wang, J., Gong, B., and Yu, S. X · 2019
Cited alongside, same era.
Xcit: Cross-covariance image transformers
Ali, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al · 2021
Later among the works it cites.
Semi-supervised learning of visual features by non-parametrically predicting view assignments with support samples
Assran, M., Caron, M., Misra, I., Bojanowski, P., Joulin, A., Ballas, N., and Rabbat, M · 2021
Later among the works it cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Later among the works it cites.
Unsupervised representation transfer for small networks: I believe i can distill on-the-fly
Choi, H. M., Kang, H., and Oh, D · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Cited alongside, same era.
Pytorch image models
Wightman, R · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Cited alongside, same era.
Compress: Self-supervised learning by compressing representations
Abbasi Koohpayegani, S., Tejankar, A., and Pirsiavash, H · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Cited alongside, same era.
Levit: A vision transformer in convnet’s clothing for faster inference
Graham, B., El-Nouby, A., Touvron, H., Stock, P., Joulin, A., Jégou, H., and Douze, M · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
Rethinking spatial dimensions of vision transformers
Heo, B., Yun, S., Han, D., Chun, S., Choe, J., and Oh, S. J · 2021
Later among the works it cites.
Benchmarking detection transfer learning with vision transformers
Li, Y., Xie, S., Chen, X., Dollar, P., He, K., and Girshick, R · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A · 2021
Later among the works it cites.
Imagenet-21k pretraining for the masses
Ridnik, T., Ben-Baruch, E., Noy, A., and Zelnik-Manor, L · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2021
Later among the works it cites.
Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wang, W., Bao, H., Huang, S., Dong, L., and Wei, F · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Wightman, R., Touvron, H., and Jégou, H · 2021
Later among the works it cites.
Exploring plain vision transformer backbones for object detection
Li, Y., Mao, H., Girshick, R., and He, K · 2022
Closest in time.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Closest in time.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
Mehta, S. and Rastegari, M · 2022
Closest in time.
Edgevits: Competing light-weight cnns on mobile devices with vision transformers
Pan, J., Bulat, A., Tan, F., Zhu, X., Dudziak, L., Li, H., Tzimiropoulos, G., and Martinez, B · 2022
Closest in time.
Touvron, H., Cord, M., and Jégou, H · 2022
Closest in time.
Simmim: A simple framework for masked image modeling
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H · 2022
Closest in time.
ibot: Image bert pre-training with online tokenizer
Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T · 2022
Closest in time.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I.-S., and Xie, S · 2023
Closest in time.