Fetching the paper…
Reading the bibliography…
Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs).
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R. 2010 · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2020 · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Tropp, J. A. 2012 · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B.; Tomioka, R.; and Srebro, N. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Regularized M-estimators with nonconvexity: Statistical and algorithmic theory for local optima
Loh, P.-L.; and Wainwright, M. J. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; and Rabinovich, A. 2015 · 2015
Earlier work this paper cites.
Scalable person re-identification: A benchmark
Zheng, L.; Shen, L.; Tian, L.; Wang, S.; Wang, J.; and Tian, Q. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S.; Girshick, R.; Dollár, P.; Tu, Z.; and He, K. 2017 · 2017
Earlier work this paper cites.
Squeeze-and-excitation networks
Hu, J.; Shen, L.; and Sun, G. 2018 · 2018
Earlier work this paper cites.
Person transfer gan to bridge domain gap for person re-identification
Wei, L.; Zhang, S.; Gao, W.; and Tian, Q. 2018 · 2018
Earlier work this paper cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Zhang, X.; Zhou, X.; Lin, M.; and Sun, J. 2018 · 2018
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Savarese, P.; Evron, I.; Soudry, D.; and Srebro, N. 2019 · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M.; and Le, Q. 2019 · 2019
Cited alongside, same era.
Sharpness-aware Minimization for Efficiently Improving Generalization
Foret, P.; Kleiner, A.; Mobahi, H.; and Neyshabur, B. 2020 · 2020
Cited alongside, same era.
How useful is self-supervised pretraining for visual tasks?
Newell, A.; and Deng, J. 2020 · 2020
Cited alongside, same era.
High-performance large-scale image recognition without normalization
CMT: Convolutional Neural Networks Meet Vision Transformers
Guo, J.; Han, K.; Wu, H.; Xu, C.; Tang, Y.; Xu, C.; and Wang, Y. 2021 · 2021
Closest in time.
Han, K.; Xiao, A.; Wu, E.; Guo, J.; Xu, C.; and Wang, Y. 2021 · 2021
Closest in time.
Transreid: Transformer-based object re-identification
He, S.; Luo, H.; Wang, P.; Wang, F.; Li, H.; and Jiang, W. 2021 · 2021
Closest in time.
Rethinking spatial dimensions of vision transformers
Heo, B.; Yun, S.; Han, D.; Chun, S.; Choe, J.; and Oh, S. J. 2021 · 2021
Closest in time.
Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brock, A.; De, S.; Smith, S. L.; and Simonyan, K. 2021 · 2021
Cited alongside, same era.
Emerging Properties in Self-Supervised Vision Transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Cited alongside, same era.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chen, C.-F.; Fan, Q.; and Panda, R. 2021 · 2021
Cited alongside, same era.
When Vision Transformers Outperform ResNets without Pretraining or Strong Data Augmentations
Chen, X.; Hsieh, C.-J.; and Gong, B. 2021 · 2021
Cited alongside, same era.
ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases
d’Ascoli, S.; Touvron, H.; Leavitt, M.; Morcos, A.; Biroli, G.; and Sagun, L. 2021 · 2021
Cited alongside, same era.
XCiT: Cross-Covariance Image Transformers
El-Nouby, A.; Touvron, H.; Caron, M.; Bojanowski, P.; Douze, M.; Joulin, A.; Laptev, I.; Neverova, N.; Synnaeve, G.; Verbeek, J.; et al. 2021 · 2021
Cited alongside, same era.
Ergen, T.; Sahiner, A.; Ozturkler, B.; Pauly, J.; Mardani, M.; and Pilanci, M. 2021 · 2021
Cited alongside, same era.
Huang, Z.; Ben, Y.; Luo, G.; Cheng, P.; Yu, G.; and Fu, B. 2021 · 2021
Closest in time.
Token Labeling: Training a 85.5% Top-1 Accuracy Vision Transformer with 56M Parameters on ImageNet
Jiang, Z.; Hou, Q.; Yuan, L.; Zhou, D.; Jin, X.; Wang, A.; and Feng, J. 2021 · 2021
Closest in time.
FoveaTer: Foveated Transformer for Image Classification
Jonnalagadda, A.; Wang, W.; and Eckstein, M. P. 2021 · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Closest in time.
Vision transformers for dense prediction
Ranftl, R.; Bochkovskiy, A.; and Koltun, V. 2021 · 2021
Closest in time.
DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Closest in time.
Going deeper with Image Transformers
Touvron, H.; Cord, M.; Sablayrolles, A.; Synnaeve, G.; and Jégou, H. 2021 · 2021
Closest in time.
Cvt: Introducing convolutions to vision transformers
Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; and Zhang, L. 2021 · 2021
Closest in time.
Early Convolutions Help Transformers See Better
Xiao, T.; Singh, M.; Mintun, E.; Darrell, T.; Dollár, P.; and Girshick, R. 2021 · 2021
Closest in time.
So-ViT: Mind Visual Tokens for Vision Transformer
Xie, J.; Zeng, R.; Wang, Q.; Zhou, Z.; and Li, P. 2021 · 2021
Closest in time.
Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer
Xu, Y.; Zhang, Z.; Zhang, M.; Sheng, K.; Li, K.; Dong, W.; Zhang, L.; Xu, C.; and Sun, X. 2021 · 2021
Closest in time.
Glance-and-Gaze Vision Transformer
Yu, Q.; Xia, Y.; Bai, Y.; Lu, Y.; Yuille, A.; and Shen, W. 2021 · 2021
Closest in time.