Fetching the paper…
Reading the bibliography…
Transformer models are revolutionizing machine learning, but their inner workings remain mysterious.
Principal Components in Regression Analysis
I. T. Jolliffe · 1904
Earlier work this paper cites.
On the relationship between self-attention and convolutional layers
J.-B. Cordonnier, A. Loukas, and M. Jaggi · 1911
Earlier work this paper cites.
Selected studies of the principle of relative frequency in language
G. K. Zipf · 1932
Earlier work this paper cites.
Visual interrogation of attention-based models for natural language inference and machine comprehension
S. Liu, T. Li, Z. Li, V. Srikumar, V. Pascucci, and P.-T. Bremer · 2007
Earlier work this paper cites.
Visualizing data using t-sne
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2010
Earlier work this paper cites.
Universal principles of design, revised and updated: 125 ways to enhance usability, influence perception, increase appeal, make better design decisions, and teach through design
W. Lidwell, K. Holden, and J. Butler · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Embedding projector: Interactive visualization and interpretation of embeddings
D. Smilkov, N. Thorat, C. Nicholson, E. Reif, F. B. Viégas, and M. Wattenberg · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam · 2017
Earlier work this paper cites.
Feature visualization
C. Olah, A. Mordvintsev, and L. Schubert · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Visual analytics in deep learning: An interrogative survey for the next frontiers
F. Hohman, M. Kahng, R. Pienta, and D. H. Chau · 2018
Earlier work this paper cites.
Embeddingvis: A visual analytics approach to comparative network embedding inspection
Q. Li, K. S. Njotoprawiro, H. Haleem, Q. Chen, C. Yi, and X. Ma · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
L. McInnes, J. Healy, and J. Melville · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Seq2seq-vis: A visual debugging tool for sequence-to-sequence models
H. Strobelt, S. Gehrmann, M. Behrisch, A. Perer, H. Pfister, and A. M. Rush · 2018
Earlier work this paper cites.
Tensor2tensor for neural machine translation
A. Vaswani, S. Bengio, E. Brevdo, F. Chollet, A. N. Gomez, S. Gouws, L. Jones, Ł. Kaiser, N. Kalchbrenner, N. Parmar, et al · 2018
Earlier work this paper cites.
A comparison of word embeddings for the biomedical natural language processing
Y. Wang, S. Liu, N. Afzal, M. Rastegar-Mojarad, L. Wang, F. Shen, P. Kingsbury, and H. Liu · 2018
Earlier work this paper cites.
Activation atlas
S. Carter, Z. Armstrong, L. Schubert, I. Johnson, and C. Olah · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Cited alongside, same era.
Summit: Scaling deep learning interpretability by visualizing activation and attribution summarizations
F. Hohman, H. Park, C. Robinson, and D. H. P. Chau · 2019
Cited alongside, same era.
Revealing the dark secrets of BERT
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky · 2019
Cited alongside, same era.
Open sesame: Getting inside BERT’s linguistic knowledge
Y. Lin, Y. C. Tan, and R. Frank · 2019
Cited alongside, same era.
Sanvis: Visual analytics for understanding self-attention networks
C. Park, I. Na, Y. Jo, S. Shin, J. Yoo, B. C. Kwon, J. Zhao, H. Noh, Y. Lee, and J. Choo · 2019
Cited alongside, same era.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, et al · 2021
Later among the works it cites.
Visqa: X-raying vision and language reasoning in transformers
T. Jaunet, C. Kervadec, R. Vuillemot, G. Antipov, M. Baccouche, and C. Wolf · 2021
Later among the works it cites.
Usevis: Visual analytics of attention-based neural embedding in information retrieval
X. Ji, Y. Tu, W. He, J. Wang, H.-W. Shen, and P.-Y. Yen · 2021
Later among the works it cites.
Intriguing properties of vision transformers
M. M. Naseer, K. Ranasinghe, S. H. Khan, M. Hayat, F. Shahbaz Khan, and M.-H. Yang · 2021
Later among the works it cites.
Dodrio: Exploring transformer models with interactive visualization
Z. J. Wang, R. Turko, and D. H. Chau · 2021
Later among the works it cites.
Transpose: Keypoint localization via transformer
S. Yang, Z. Quan, M. Nie, and W. Yang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
E. Reif, A. Yuan, M. Wattenberg, F. B. Viegas, A. Coenen, A. Pearce, and B. Kim · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Cited alongside, same era.
A multiscale visualization of attention in the transformer model
J. Vig · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model
J. Vig and Y. Belinkov · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov · 2019
Cited alongside, same era.
Later among the works it cites.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
E. Aflalo, M. Du, S.-Y. Tseng, Y. Liu, C. Wu, N. Duan, and V. Lal · 2022
Later among the works it cites.
Embedding comparator: Visualizing differences in global structure and local neighborhoods via small multiples
A. Boggust, B. Carter, and A. Satyanarayan · 2022
Later among the works it cites.
Toy models of superposition
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, et al · 2022
Later among the works it cites.
What do vision transformers learn? a visual exploration
A. Ghiasi, H. Kazemi, E. Borgnia, S. Reich, M. Shu, M. Goldblum, A. G. Wilson, and T. Goldstein · 2022
Later among the works it cites.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Later among the works it cites.
How do vision transformers work?
N. Park and S. Kim · 2022
Later among the works it cites.
Visual comparison of language model adaptation
R. Sevastjanova, E. Cakmak, S. Ravfogel, R. Cotterell, and M. El-Assady · 2022
Later among the works it cites.
Emblaze: Illuminating machine learning representations through interactive comparison of embedding spaces
V. Sivaraman, Y. Wu, and A. Perer · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Later among the works it cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus · 2022
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Closest in time.
Visualizing and understanding patch interactions in vision transformer
J. Ma, Y. Bai, B. Zhong, W. Zhang, T. Yao, and T. Mei · 2023
Closest in time.
Self-attention in vision transformers performs perceptual grouping, not attention
P. Mehrani and J. K. Tsotsos · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Closest in time.