Fetching the paper…
Reading the bibliography…
By computing the rank correlation between attention weights and feature-additive explanation methods, previous analyses either invalidate or support the role of attention-based explanations as a faithful and plausible measure of salience.
Fine-grained sentiment analysis with faithful attention
Zhong, R., Shao, S., and McKeown, K. R · 1908
Earlier work this paper cites.
Can I trust the explainer? verifying post-hoc explanatory methods
Camburu, O., Giunchiglia, E., Foerster, J., Lukasiewicz, T., and Blunsom, P · 1910
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 1910
Earlier work this paper cites.
Understanding multi-head attention in abstractive summarization
Baan, J., ter Hoeve, M., van der Wees, M., Schuth, A., and de Rijke, M · 1911
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Timeshap: Explaining recurrent models through sequence perturbations
Bento, J., Saleiro, P., Cruz, A. F., Figueiredo, M. A. T., and Bizarro, P · 2012
Earlier work this paper cites.
Parsing with compositional vector grammars
Socher, R., Bauer, J., Manning, C. D., and Ng, A. Y · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Caruana, R., Lou, Y., Gehrke, J., Koch, P., Sturm, M., and Elhadad, N · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Interpretation of prediction models using the input gradient
Hechtlinger, Y · 2016
Earlier work this paper cites.
Investigating the influence of noise and distractors on the interpretation of neural networks
Kindermans, P., Schütt, K., Müller, K., and Dähne, S · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Li, J., Monroe, W., and Jurafsky, D · 2016
Earlier work this paper cites.
”why should i trust you?”: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
Algorithms in the criminal justice system: Assessing the use of risk assessments in sentencing
Kehl, D. and Kessler, S. A · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
XNLI: evaluating cross-lingual sentence representations
Conneau, A., Rinott, R., Lample, G., Williams, A., Bowman, S. R., Schwenk, H., and Stoyanov, V · 2018
Cited alongside, same era.
On the convergence proof of amsgrad and a new version
Tran, P. T. and Phong, L. T · 2019
Later among the works it cites.
Attention is not not explanation
Wiegreffe, S. and Pinter, Y · 2019
Later among the works it cites.
Quantifying attention flow in transformers
Abnar, S. and Zuidema, W · 2020
Later among the works it cites.
A diagnostic study of explainability techniques for text classification
Atanasova, P., Simonsen, J. G., Lioma, C., and Augenstein, I · 2020
Later among the works it cites.
The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
Bastings, J. and Filippova, K · 2020
Later among the works it cites.
On identifiability in transformers
Brunner, G., Liu, Y., Pascual, D., Richter, O., Ciaramita, M., and Wattenhofer, R · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Allennlp: A deep semantic natural language processing platform
Gardner, M., Grus, J., Neumann, M., Tafjord, O., Dasigi, P., Liu, N. F., Peters, M. E., Schmitz, M., and Zettlemoyer, L · 2018
Cited alongside, same era.
Interpretable credit application predictions with counterfactual explanations
McGrath, R., Costabello, L., Van, C. L., Sweeney, P., Kamiab, F., Shen, Z., and Lécué, F · 2018
Cited alongside, same era.
Perturbation-Based Explanations of Prediction Models , pp. 159–175
Robnik-Šikonja, M. and Bohanec, M · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Cited alongside, same era.
Adaptively sparse transformers
Correia, G. M., Niculae, V., and Martins, A. F. T · 2019
Cited alongside, same era.
ERASER: A benchmark to evaluate rationalized NLP models
DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., and Wallace, B. C · 2020
Later among the works it cites.
Why attention is not explanation: Surgical intervention and causal reasoning about neural models
Grimsley, C., Mayfield, E., and R.S. Bursten, J · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Jacovi, A. and Goldberg, Y · 2020
Later among the works it cites.
Attention is not only a weight: Analyzing transformers with vector norms
Kobayashi, G., Kuribayashi, T., Yokoi, S., and Inui, K · 2020
Later among the works it cites.
Towards transparent and explainable attention models
Mohankumar, A. K., Nema, P., Narasimhan, S., Khapra, M. M., Srinivasan, B. V., and Ravindran, B · 2020
Later among the works it cites.
Understanding attention for text classification
Sun, X. and Lu, W · 2020
Later among the works it cites.
Staying true to your word: (how) can attention become explanation?
Tutek, M. and Snajder, J · 2020
Later among the works it cites.
Attention flows are shapley value explanations
Ethayarajh, K. and Jurafsky, D · 2021
Closest in time.
Is sparse attention more interpretable?
Meister, C., Lazov, S., Augenstein, I., and Cotterell, R · 2021
Closest in time.
Teach me to explain: A review of datasets for explainable NLP
Wiegreffe, S. and Marasovic, A · 2021
Closest in time.
Evaluating the correctness of explainable AI algorithms for classification
Yalcin, O., Fan, X., and Liu, S · 2021
Closest in time.