Fetching the paper…
Reading the bibliography…
While recent years have witnessed the emergence of various explainable methods in machine learning, to what degree the explanations really represent the reasoning process behind the model prediction -- namely, the faithfulness of explanation -- is still an open problem.
THE TREATMENT OF TIES IN RANKING PROBLEMS
M. G. KENDALL. 1945 · 1945
Earlier work this paper cites.
CRC standard probability and statistics tables and formulae
Daniel Zwillinger and Stephen Kokoska. 1999 · 1999
Earlier work this paper cites.
Thumbs up? sentiment classification using machine learning techniques
Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002 · 2002
Earlier work this paper cites.
Using “annotator rationales” to improve machine learning for text categorization
Omar Zaidan, Jason Eisner, and Christine Piatko. 2007 · 2007
Earlier work this paper cites.
Explaining classifications for individual instances
Marko Robnik-Šikonja and Igor Kononenko. 2008 · 2008
Earlier work this paper cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John Dickerson, and Keegan Hines. 2020 · 2010
Earlier work this paper cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C Lipton, Graham Neubig, and William W Cohen. 2020 · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
David Alvarez-Melis and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Cited alongside, same era.
UCI machine learning repository
Dheeru Dua and Casey Graff. 2017 · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2017 · 2017
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Rethinking cooperative rationalization: Introspective extraction and complement control
Mo Yu, Shiyu Chang, Yang Zhang, and Tommi Jaakkola. 2019 · 2019
Later among the works it cites.
The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
Jasmijn Bastings and Katja Filippova. 2020 · 2020
Later among the works it cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Later among the works it cites.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. 2018 · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Cited alongside, same era.
Comparing automatic and human evaluation of local explanations for text classification
Dong Nguyen. 2018 · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
The promise and peril of human evaluation for model interpretability
Bernease Herman. 2019 · 2019
Cited alongside, same era.
Peter Hase and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Interpretable Machine Learning
C. Molnar. 2020 · 2020
Later among the works it cites.
Explaining machine learning classifiers through diverse counterfactual explanations
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan. 2020 · 2020
Later among the works it cites.
Variable instance-level explainability for text classification
George Chrysostomou and Nikolaos Aletras. 2021 · 2021
Closest in time.
Evaluating saliency methods for neural language models
Shuoyang Ding and Philipp Koehn. 2021 · 2021
Closest in time.
Evaluating explanations for reading comprehension with realistic counterfactuals
Xi Ye, Rohan Nair, and Greg Durrett. 2021 · 2021
Closest in time.