Fetching the paper…
Reading the bibliography…
Despite the growing use of transformer models in computer vision, a mechanistic understanding of these networks is still needed.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
On identifiability in transformers
G. Brunner, Y. Liu, D. Pascual, O. Richter, M. Ciaramita, and R. Wattenhofer · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
A. Ravichander, Y. Belinkov, and E. Hovy · 2020
Earlier work this paper cites.
Understanding robustness of transformers for image classification
S. Bhojanapalli, A. Chakrabarti, D. Glasner, D. Li, T. Unterthiner, and A. Veit · 2021
Earlier work this paper cites.
Overinterpretation reveals image classification model pathologies
B. Carter, S. Jain, J. W. Mueller, and D. Gifford · 2021
Earlier work this paper cites.
Transformer interpretability beyond attention visualization
H. Chefer, S. Gur, and L. Wolf · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
M. Geva, R. Schuster, J. Berant, and O. Levy · 2021
Cited alongside, same era.
Intriguing properties of vision transformers
M. M. Naseer, K. Ranasinghe, S. H. Khan, M. Hayat, F. Shahbaz Khan, and M.-H. Yang · 2021
Cited alongside, same era.
Do vision transformers see like convolutional neural networks?
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy · 2021
Cited alongside, same era.
Imagenet-21k pretraining for the masses
T. Ridnik, E. Ben-Baruch, A. Noy, and L. Zelnik-Manor · 2021
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
M. Geva, A. Caciularu, K. Wang, and Y. Goldberg · 2022
Later among the works it cites.
What do vision transformers learn? a visual exploration
A. Ghiasi, H. Kazemi, E. Borgnia, S. Reich, M. Shu, M. Goldblum, A. G. Wilson, and T. Goldstein · 2022
Later among the works it cites.
On improving adversarial transferability of vision transformers
M. Naseer, K. Ranasinghe, S. Khan, F. Khan, and F. Porikli · 2022
Later among the works it cites.
How do vision transformers work?
N. Park and S. Kim · 2022
Later among the works it cites.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Dar, M. Geva, A. Gupta, and J. Berant · 2022
Cited alongside, same era.
Large-scale unsupervised semantic segmentation
S. Gao, Z.-Y. Li, M.-H. Yang, M.-M. Cheng, J. Han, and P. Torr · 2022
Cited alongside, same era.
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt · 2023
Closest in time.
Jump to conclusions: Short-cutting transformers with linear transformations
A. Y. Din, T. Karidi, L. Choshen, and M. Geva · 2023
Closest in time.