Fetching the paper…
Reading the bibliography…
Vision Transformer(ViT) is one of the most widely used models in the computer vision field with its great performance on various tasks.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E.; Talbot, D.; Moiseev, F.; Sennrich, R.; and Titov, I. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y.; Boser, B.; Denker, J. S.; Henderson, D.; Howard, R. E.; Hubbard, W.; and Jackel, L. D. 1989 · 1989
Earlier work this paper cites.
Quantifying attention flow in transformers
Abnar, S.; and Zuidema, W. 2020 · 2005
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Draelos, R. L.; and Carin, L. 2020 · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
Wah, C.; Branson, S.; Welinder, P.; Perona, P.; and Belongie, S. 2011 · 2011
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
Everingham, M.; Van Gool, L.; Williams, C. K. I.; Winn, J.; and Zisserman, A. 2012 · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S.; Binder, A.; Montavon, G.; Klauschen, F.; Müller, K.-R.; and Samek, W. 2015 · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; and Rabinovich, A. 2015 · 2015
Earlier work this paper cites.
Layer-wise relevance propagation for deep neural network architectures
Binder, A.; Bach, S.; Montavon, G.; Müller, K.-R.; and Samek, W. 2016a · 2016
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Binder, A.; Montavon, G.; Lapuschkin, S.; Müller, K.-R.; and Samek, W. 2016b · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Cited alongside, same era.
Evaluating the visualization of what a deep neural network has learned
Samek, W.; Binder, A.; Montavon, G.; Lapuschkin, S.; and Müller, K.-R. 2016 · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Zhou, B.; Khosla, A.; Lapedriza, A.; Oliva, A.; and Torralba, A. 2016 · 2016
Cited alongside, same era.
Explaining nonlinear classification decisions with deep taylor decomposition
Combinational class activation maps for weakly supervised object localization
Yang, S.; Kim, Y.; Kim, Y.; and Kim, C. 2020 · 2020
Later among the works it cites.
Transformer interpretability beyond attention visualization
Chefer, H.; Gur, S.; and Wolf, L. 2021 · 2021
Later among the works it cites.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chen, C.-F. R.; Fan, Q.; and Panda, R. 2021 · 2021
Later among the works it cites.
Localvit: Bringing locality to vision transformers
Li, Y.; Zhang, K.; Cao, J.; Timofte, R.; and Van Gool, L. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Montavon, G.; Lapuschkin, S.; Binder, A.; Samek, W.; and Müller, K.-R. 2017 · 2017
Cited alongside, same era.
Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks
Chattopadhay, A.; Sarkar, A.; Howlader, P.; and Balasubramanian, V. N. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Cited alongside, same era.
PyTorch Image Models
Wightman, R. 2019 · 2019
Cited alongside, same era.
Naseer, M. M.; Ranasinghe, K.; Khan, S. H.; Hayat, M.; Shahbaz Khan, F.; and Yang, M.-H. 2021 · 2021
Later among the works it cites.
Informative Class Activation Maps
Qin, Z.; Kim, D.; and Gedeon, T. 2021 · 2021
Later among the works it cites.
Vision transformers for dense prediction
Ranftl, R.; Bochkovskiy, A.; and Koltun, V. 2021 · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Later among the works it cites.
Are convolutional neural networks or transformers more like human vision?
Tuli, S.; Dasgupta, I.; Grant, E.; and Griffiths, T. L. 2021 · 2021
Later among the works it cites.
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolutions
Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; and Shao, L. 2021 · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. 2021 · 2021
Later among the works it cites.