Fetching the paper…
Reading the bibliography…
The field of natural language processing has reached breakthroughs with the advent of transformers.
What does BERT look at? an analysis of bert’s attention,
K. Clark, U. Khandelwal, O. Levy, C. D. Manning, · 1906
Earlier work this paper cites.
Augmenting self-attention with persistent memory,
S. Sukhbaatar, E. Grave, G. Lample, H. Jégou, A. Joulin, · 1907
Earlier work this paper cites.
One explanation does not fit all: A toolkit and taxonomy of AI explainability techniques,
V. Arya, R. K. E. Bellamy, P. Chen, A. Dhurandhar, M. Hind, S. C. Hoffman, S. Houde, Q. V. Liao, R. Luss, A. Mojsilovic, S. Mourad, P. Pedemonte, R. Raghavendra, J. T. Richards, P. Sattigeri, K. Shanmugam, M. Singh, K. R. Varshney, D. Wei, Y. Zhang, · 1909
Earlier work this paper cites.
On the linguistic representational power of neural machine translation models,
Y. Belinkov, N. Durrani, F. Dalvi, H. Sajjad, J. R. Glass, · 1911
Earlier work this paper cites.
2004
Earlier work this paper cites.
Opportunities and challenges in explainable artificial intelligence (XAI): A survey,
A. Das, P. Rad, · 2006
Earlier work this paper cites.
Transformer with depth-wise LSTM,
H. Xu, Q. Liu, D. Xiong, J. van Genabith, · 2007
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models,
N. Durrani, H. Sajjad, F. Dalvi, Y. Belinkov, · 2010
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories,
M. Geva, R. Schuster, J. Berant, O. Levy, · 2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability,
B. Kim, R. Khanna, O. O. Koyejo, · 2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, · 2017
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences,
T. Miller, · 2017
Earlier work this paper cites.
D. Hupkes, S. Veldhoen, W. Zuidema, · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks,
M. Sundararajan, A. Taly, Q. Yan, · 2017
Earlier work this paper cites.
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, R. Sayres, · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M. Chang, K. Lee, K. Toutanova, · 2018
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (xai),
A. Adadi, M. Berrada, · 2018
Cited alongside, same era.
A survey of methods for explaining black box models,
R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, D. Pedreschi, · 2018
Cited alongside, same era.
Explaining explanations: An approach to evaluating interpretability of machine learning,
L. H. Gilpin, D. Bau, B. Z. Yuan, A. Bajwa, M. A. Specter, L. Kagal, · 2018
Cited alongside, same era.
Explainable prediction of medical codes from clinical text,
J. Mullenbach, S. Wiegreffe, J. Duke, J. Sun, J. Eisenstein, · 2018
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties,
A. Conneau, G. Kruszewski, G. Lample, L. Barrault, M. Baroni, · 2018
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline,
I. Tenney, D. Das, E. Pavlick, · 2019
Later among the works it cites.
Improving transformer models by reordering their sublayers,
O. Press, N. A. Smith, O. Levy, · 2020
Later among the works it cites.
A survey of the state of explainable AI for natural language processing,
M. Danilevsky, K. Qian, R. Aharonov, Y. Katsis, B. Kawas, P. Sen, · 2020
Later among the works it cites.
What happens to BERT embeddings during fine-tuning?,
A. Merchant, E. Rahimtoroghi, E. Pavlick, I. Tenney, · 2020
Later among the works it cites.
On the Interplay Between Fine-tuning and Sentence-level Probing for Linguistic Knowledge in Pre-trained Transformers,
M. Mosbach, A. Khokhlova, M. A. Hedderich, D. Klakow, · 2020
Later among the works it cites.
Unsupervised domain clusters in pretrained language models,
R. Aharoni, Y. Goldberg, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Analysis methods in neural language processing: A survey,
Y. Belinkov, J. R. Glass, · 2018
Cited alongside, same era.
Seq2seq-vis: A visual debugging tool for sequence-to-sequence models,
H. Strobelt, S. Gehrmann, M. Behrisch, A. Perer, H. Pfister, A. M. Rush, · 2018
Cited alongside, same era.
Deep contextualized word representations,
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, L. Zettlemoyer, · 2018
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks,
D. Alvarez Melis, T. Jaakkola, · 2018
Cited alongside, same era.
Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning,
M. Nasr, R. Shokri, A. Houmansadr, · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model,
J. Vig, Y. Belinkov, · 2019
Cited alongside, same era.
Neurox: A toolkit for analyzing individual neurons in neural networks,
F. Dalvi, A. Nortonsmith, D. A. Bau, Y. Belinkov, H. Sajjad, N. Durrani, J. Glass, · 2019
Cited alongside, same era.
Later among the works it cites.
2021
Later among the works it cites.
Knowledge neurons in pretrained transformers,
D. Dai, L. Dong, Y. Hao, Z. Sui, F. Wei, · 2021
Later among the works it cites.
Of non-linearity and commutativity in BERT,
S. Zhao, D. Pascual, G. Brunner, R. Wattenhofer, · 2021
Later among the works it cites.
How transfer learning impacts linguistic knowledge in deep NLP models?,
N. Durrani, H. Sajjad, F. Dalvi, · 2021
Later among the works it cites.
Ecco: An open source library for the explainability of transformer language models,
J. Alammar, · 2021
Later among the works it cites.
An interpretability illusion for BERT,
T. Bolukbasi, A. Pearce, A. Yuan, A. Coenen, E. Reif, F. B. Viégas, M. Wattenberg, · 2021
Later among the works it cites.
Explainable ai: A review of machine learning interpretability methods,
P. Linardatos, V. Papastefanopoulos, S. Kotsiantis, · 2021
Later among the works it cites.
Neuron-level interpretation of deep NLP models: A survey,
H. Sajjad, N. Durrani, F. Dalvi, · 2021
Later among the works it cites.
Best of both worlds: local and global explanations with human-understandable concepts,
J. Schrouff, S. Baur, S. Hou, D. Mincu, E. Loreaux, R. Blanes, J. Wexler, A. Karthikesalingam, B. Kim, · 2021
Later among the works it cites.
Explainable artificial intelligence: An updated perspective,
A. Krajna, M. Kovac, M. Brcic, A. Šarčević, · 2022
Later among the works it cites.
2022
Later among the works it cites.