Fetching the paper…
Reading the bibliography…
Attention-based architectures, in particular transformers, are at the heart of a technological revolution.
A value for n n -person games
Shapley, L. S · 1953
Earlier work this paper cites.
On context-free languages and push-down automata
Schützenberger, M. P · 1963
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I · 2003
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Extraction of salient sentences from labelled documents
Denil, M., Demiraj, A., and De Freitas, N · 2014
Earlier work this paper cites.
Sak, H., Senior, A., and Beaufays, F · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K. H., and Bengio, Y · 2015
Earlier work this paper cites.
Bidirectional lstm-crf models for sequence tagging
Huang, Z., Xu, W., and Yu, K · 2015
Earlier work this paper cites.
Visualizing and understanding neural models in nlp
Li, J., Chen, X., Hovy, E., and Jurafsky, D · 2016
Earlier work this paper cites.
“Why should I trust you?” Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Earlier work this paper cites.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
Poerner, N., Schütze, H., and Roth, B · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Ribeiro, M. T., Singh, S., and Guestrin, C · 2018
Earlier work this paper cites.
Evaluating recurrent neural network explanations
Arras, L., Osman, A., Müller, K.-R., and Samek, W · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Attention is not explanation
Jain, S. and Wallace, B. C · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Is attention interpretable?
Serrano, S. and Smith, N. A · 2019
Cited alongside, same era.
Generating token-level explanations for natural language inference
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model
Vig, J. and Belinkov, Y · 2019
Cited alongside, same era.
Order in the court: Explainable AI methods prone to disagreement
Neely, M., Schouten, S. F., Bleeker, M. J. R., and Lucic, A · 2021
Later among the works it cites.
Effective attention sheds light on interpretability
Sun, K. and Marasović, A · 2021
Later among the works it cites.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Later among the works it cites.
Is attention explanation? an introduction to the debate
Bibal, A., Cardon, R., Alfter, D., Wilkens, R., Wang, X., François, T., and Watrin, P · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L., Goel, S., Kakade, S., and Zhang, C · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M., and Li, Y · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wiegreffe, S. and Pinter, Y · 2019
Cited alongside, same era.
A diagnostic study of explainability techniques for text classification
Atanasova, P., Simonsen, J. G., Lioma, C., and Augenstein, I · 2020
Cited alongside, same era.
The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
Bastings, J. and Filippova, K · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
On identifiability in transformers
Brunner, G., Liu, Y., Pascual, D., Richter, O., Ciaramita, M., and Wattenhofer, R · 2020
Cited alongside, same era.
Attention in natural language processing
Galassi, A., Lippi, M., and Torroni, P · 2020
Cited alongside, same era.
Attention is not only a weight: Analyzing transformers with vector norms
Kobayashi, G., Kuribayashi, T., Yokoi, S., and Inui, K · 2020
Cited alongside, same era.
Later among the works it cites.
Delivering trustworthy ai through formal xai
Marques-Silva, J. and Ignatiev, A · 2022
Later among the works it cites.
A Song of (Dis)agreement: Evaluating the Evaluation of Explainable Artificial Intelligence in Natural Language Processing
Neely, M., Schouten, S. F., Bleeker, M., and Lucic, A · 2022
Later among the works it cites.
Formal algorithms for transformers
Phuong, M. and Hutter, M · 2022
Later among the works it cites.
Benchmarking and survey of explanation methods for black box models
Bodria, F., Giannotti, F., Guidotti, R., Naretto, F., Pedreschi, D., and Rinzivillo, S · 2023
Later among the works it cites.
What can a Single Attention Layer Learn? A Study Through the Random Features Lens
Fu, H., Guo, T., Bai, Y., and Mei, S · 2023
Later among the works it cites.
How do transformers learn topic structure: Towards a mechanistic understanding
Li, Y., Li, Y., and Risteski, A · 2023
Later among the works it cites.
An attention matrix for every decision: Faithfulness-based arbitration among multiple attention-based interpretations of transformers in text classification
Mylonas, N., Mollas, I., and Tsoumakas, G · 2023
Later among the works it cites.
Transformers as support vector machines
Tarzanagh, D. A., Li, Y., Thrampoulidis, C., and Oymak, S · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Cui, H., Behrens, F., Krzakala, F., and Zdeborová, L · 2024
Closest in time.
Attention with Markov: A framework for principled analysis of transformers via markov chains
Makkuva, A. V., Bondaschi, M., Girish, A., Nagle, A., Jaggi, M., Kim, H., and Gastpar, M · 2024
Closest in time.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
Wen, K., Li, Y., Liu, B., and Risteski, A · 2024
Closest in time.