Fetching the paper…
Reading the bibliography…
A common approach to quantifying neural text classifier interpretability is to calculate faithfulness metrics based on iteratively masking salient input tokens and measuring changes in the model prediction.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Jacovi, A.; and Goldberg, Y. 2020 · 2004
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
Explaining Predictions of Non-Linear Classifiers in NLP
Arras, L.; Horn, F.; Montavon, G.; Müller, K.-R.; and Samek, W. 2016 · 2016
Earlier work this paper cites.
” Why should i trust you?” Explaining the predictions of any classifier
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Lundberg, S. M.; and Lee, S.-I. 2017 · 2017
Earlier work this paper cites.
Evaluating the Visualization of What a Deep Neural Network Has Learned
Samek, W.; Binder, A.; Montavon, G.; Lapuschkin, S.; and Müller, K.-R. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M.; Taly, A.; and Yan, Q. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Generating Natural Language Adversarial Examples
Alzantot, M.; Sharma, Y.; Elgohary, A.; Ho, B.-J.; Srivastava, M.; and Chang, K.-W. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Dixon, L.; Li, J.; Sorensen, J.; Thain, N.; and Vasserman, L. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-Box Adversarial Examples for Text Classification
Ebrahimi, J.; Rao, A.; Lowd, D.; and Dou, D. 2018 · 2018
Cited alongside, same era.
Pathologies of Neural Models Make Interpretations Difficult
Feng, S.; Wallace, E.; Grissom II, A.; Iyyer, M.; Rodriguez, P.; and Boyd-Graber, J. 2018 · 2018
Cited alongside, same era.
Black-box generation of adversarial text sequences to evade deep learning classifiers
Gao, J.; Lanchantin, J.; Soffa, M. L.; and Qi, Y. 2018 · 2018
Cited alongside, same era.
The Mythos of Model Interpretability: In Machine Learning, the Concept of Interpretability is Both Important and Slippery
Lipton, Z. C. 2018 · 2018
Cited alongside, same era.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
McInnes, L.; and Healy, J. 2018 · 2018
Cited alongside, same era.
A Diagnostic Study of Explainability Techniques for Text Classification
Atanasova, P.; Simonsen, J. G.; Lioma, C.; and Augenstein, I. 2020 · 2020
Later among the works it cites.
ERASER: A Benchmark to Evaluate Rationalized NLP Models
DeYoung, J.; Jain, S.; Rajani, N. F.; Lehman, E.; Xiong, C.; Socher, R.; and Wallace, B. C. 2020 · 2020
Later among the works it cites.
Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
Jin, D.; Jin, Z.; Zhou, J. T.; and Szolovits, P. 2020 · 2020
Later among the works it cites.
Towards Improving Adversarial Training of NLP Models
Yoo, J. Y.; and Qi, Y. 2021 · 2021
Later among the works it cites.
On the Lack of Robust Interpretability of Neural Text Classifiers
Zafar, M. B.; Donini, M.; Slack, D.; Archambeau, C.; Das, S.; and Kenthapadi, K. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Comparing automatic and human evaluation of local explanations for text classification
Nguyen, D. 2018 · 2018
Cited alongside, same era.
BERT as Service: GitHub Repository
Xiao, H. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Jigsaw unintended bias in toxicity classification
Jigsaw. 2019 · 2019
Cited alongside, same era.
Twitter climate change sentiment dataset
Qian, E. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Cited alongside, same era.
Asthana, K.; Xie, Z.; You, W.; Noack, A.; Brophy, J.; Singh, S.; and Lowd, D. 2022 · 2022
Later among the works it cites.
Adversarial robustness of neural-statistical features in detection of generative transformers
Crothers, E.; Japkowicz, N.; Viktor, H.; and Branco, P. 2022 · 2022
Later among the works it cites.
Detecting Word-Level Adversarial Text Attacks via SHapley Additive exPlanations
Huber, L.; Kühn, M. A.; Mosca, E.; and Groh, G. 2022 · 2022
Later among the works it cites.
Fooling Explanations in Text Classifiers
Ivankay, A.; Girardi, I.; Marchiori, C.; and Frossard, P. 2022 · 2022
Later among the works it cites.
Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining
Madsen, A.; Meade, N.; Adlakha, V.; and Reddy, S. 2022 · 2022
Later among the works it cites.
Adversarial training improves model interpretability in single-cell RNA-seq analysis
Sadria, M.; Layton, A.; and Bader, G. 2023 · 2023
Closest in time.