Fetching the paper…
Reading the bibliography…
Transformers, composed of multiple self-attention layers, hold strong promises toward a generic learning primitive applicable to different data modalities, including the recent breakthroughs in computer vision achieving state-of-the-art (SOTA) standard accuracy.
Micro-Batch Training with Batch-Channel Normalization and Weight Standardization
Qiao, S.; Wang, H.; Liu, C.; Shen, W.; and Yuille, A. 2019 · 1903
Earlier work this paper cites.
Selfie: Self-supervised pretraining for image embedding
Trinh, T. H.; Luong, M.-T.; and Le, Q. V. 2019 · 1906
Earlier work this paper cites.
Instance adaptive adversarial training: Improved accuracy tradeoffs in neural nets
Balaji, Y.; Goldstein, T.; and Hoffman, J. 2019 · 1910
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Hendrycks, D.; Basart, S.; Mu, N.; Kadavath, S.; Wang, F.; Dorundo, E.; Desai, R.; Zhu, T.; Parajuli, S.; Guo, M.; Song, D.; Steinhardt, J.; and Gilmer, J. 2020 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Cao, Y.; Xu, J.; Lin, S.; Wei, F.; and Hu, H. 2020 · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2020 · 2012
Earlier work this paper cites.
Attention in natural scenes: contrast affects rapid visual processing and fixations alike
Hart, B. M.; Schmidt, H. C. E. F.; Klein-Harmeyer, I.; and Einhäuser, W. 2013 · 2013
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Zagoruyko, S.; and Komodakis, N. 2016 · 2015
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Identity Mappings in Deep Residual Networks
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G.; Sun, Y.; Liu, Z.; Sedra, D.; and Weinberger, K. Q. 2016 · 2016
Earlier work this paper cites.
Deepfool: a simple and accurate method to fool deep neural networks
Moosavi-Dezfooli, S.-M.; Fawzi, A.; and Frossard, P. 2016 · 2016
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
DeVries, T.; and Taylor, G. W. 2017 · 2017
Cited alongside, same era.
Measuring the tendency of cnns to learn surface statistical regularities
Jo, J.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Cited alongside, same era.
Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
Sun, C.; Shrivastava, A.; Singh, S.; and Gupta, A. 2017 · 2017
Cited alongside, same era.
Big Transfer (BiT): General Visual Representation Learning
Kolesnikov, A.; Beyer, L.; Zhai, X.; Puigcerver, J.; Yung, J.; Gelly, S.; and Houlsby, N. 2020 · 2020
Later among the works it cites.
Hold me tight! Influence of discriminative features on deep network boundaries
Ortiz-Jimenez, G.; Modas, A.; Moosavi, S.-M.; and Frossard, P. 2020 · 2020
Later among the works it cites.
Designing Network Design Spaces
Radosavovic, I.; Kosaraju, R.; Girshick, R.; He, K.; and Dollar, P. 2020 · 2020
Later among the works it cites.
Self-Training With Noisy Student Improves ImageNet Classification
Xie, Q.; Luong, M.-T.; Hovy, E.; and Le, Q. V. 2020 · 2020
Later among the works it cites.
Understanding Robustness of Transformers for Image Classification
Bhojanapalli, S.; Chakrabarti, A.; Glasner, D.; Li, D.; Unterthiner, T.; and Veit, A. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks
Hu, J.; Shen, L.; Albanie, S.; Sun, G.; and Vedaldi, A. 2018 · 2018
Cited alongside, same era.
Towards Deep Learning Models Resistant to Adversarial Attacks
Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; and Vladu, A. 2018 · 2018
Cited alongside, same era.
Image Transformer
Parmar, N.; Vaswani, A.; Uszkoreit, J.; Kaiser, L.; Shazeer, N.; Ku, A.; and Tran, D. 2018 · 2018
Cited alongside, same era.
Group Normalization
Wu, Y.; and He, K. 2018 · 2018
Cited alongside, same era.
Unlabeled Data Improves Adversarial Robustness
Carmon, Y.; Raghunathan, A.; Schmidt, L.; Duchi, J. C.; and Liang, P. S. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Closest in time.
When Vision Transformers Outperform ResNets without Pretraining or Strong Data Augmentations
Chen, X.; Hsieh, C.-J.; and Gong, B. 2021 · 2021
Closest in time.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Closest in time.
Sharpness-aware Minimization for Efficiently Improving Generalization
Foret, P.; Kleiner, A.; Mobahi, H.; and Neyshabur, B. 2021 · 2021
Closest in time.
Natural Adversarial Examples
Hendrycks, D.; Zhao, K.; Basart, S.; Steinhardt, J.; and Song, D. 2021 · 2021
Closest in time.
Token labeling: Training a 85.5% top-1 accuracy vision transformer with 56m parameters on imagenet
Jiang, Z.; Hou, Q.; Yuan, L.; Zhou, D.; Jin, X.; Wang, A.; and Feng, J. 2021 · 2021
Closest in time.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Closest in time.
On the Robustness of Vision Transformers to Adversarial Examples
Mahmood, K.; Mahmood, R.; and Van Dijk, M. 2021 · 2021
Closest in time.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Closest in time.
Do Vision Transformers See Like Convolutional Neural Networks?
Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021 · 2021
Closest in time.
On the Adversarial Robustness of Visual Transformers
Shao, R.; Shi, Z.; Yi, J.; Chen, P.-Y.; and Hsieh, C.-J. 2021 · 2021
Closest in time.
How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Steiner, A.; Kolesnikov, A.; Zhai, X.; Wightman, R.; Uszkoreit, J.; and Beyer, L. 2021 · 2021
Closest in time.
EfficientNetV2: Smaller Models and Faster Training
Tan, M.; and Le, Q. 2021 · 2021
Closest in time.
Are Convolutional Neural Networks or Transformers more like human vision?
Tuli, S.; Dasgupta, I.; Grant, E.; and Griffiths, T. L. 2021 · 2021
Closest in time.
Noise or Signal: The Role of Image Backgrounds in Object Recognition
Xiao, K.; Engstrom, L.; Ilyas, A.; and Madry, A. 2021 · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Tay, F. E.; Feng, J.; and Yan, S. 2021 · 2021
Closest in time.