Fetching the paper…
Reading the bibliography…
Despite the rising popularity of saliency-based explanations, the research community remains at an impasse, facing doubts concerning their purpose, efficacy, and tendency to contradict each other.
Modeling annotators: A generative approach to learning from annotator rationales
Omar Zaidan and Jason Eisner · 2008
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned, 2015
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Bach, and Klaus-Robert Müller · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
The mythos of model interpretability
Zachary C Lipton · 2018
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko · 2018
Earlier work this paper cites.
Comparing automatic and human evaluation of local explanations for text classification
Dong Nguyen · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Sofia Serrano and Noah A Smith · 2019
Cited alongside, same era.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon · 2019
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton · 2020
Later among the works it cites.
Gradient-based analysis of nlp models is manipulable
Junlin Wang, Jens Tuyls, Eric Wallace, and Sameer Singh · 2020
Later among the works it cites.
Fairwashing explanations with off-manifold detergent
Christopher Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, and Pan Kessel · 2020
Later among the works it cites.
The explanation game: Towards prediction explainability through sparse communication
Marcos V Treviso and André FT Martins · 2020
Later among the works it cites.
Have we learned to explain?: How interpretability methods can learn to encode predictions in their interpretations
Neil Jethani, Mukund Sudarshan, Yindalon Aphinyanaphongs, and Rajesh Ranganath · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Cited alongside, same era.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Cited alongside, same era.
Inferring which medical treatments work from reports of clinical trials
Eric Lehman, Jay DeYoung, Regina Barzilay, and Byron C Wallace · 2019
Cited alongside, same era.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Cited alongside, same era.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Cited alongside, same era.
Explaining by removing: A unified framework for model explanation
Ian C Covert, Scott Lundberg, and Su-In Lee · 2021
Later among the works it cites.
The out-of-distribution problem in explainability and search methods for feature importance explanations
Peter Hase, Harry Xie, and Mohit Bansal · 2021
Later among the works it cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C Lipton, Graham Neubig, and William W Cohen · 2022
Later among the works it cites.
The disagreement problem in explainable machine learning: A practitioner’s perspective
Satyapriya Krishna, Tessa Han, Alex Gu, Javin Pombra, Shahin Jabbari, Steven Wu, and Himabindu Lakkaraju · 2022
Later among the works it cites.
Openxai: Towards a transparent evaluation of model explanations
Chirag Agarwal, Satyapriya Krishna, Eshika Saxena, Martin Pawelczyk, Nari Johnson, Isha Puri, Marinka Zitnik, and Himabindu Lakkaraju · 2022
Later among the works it cites.
A comparative study of faithfulness metrics for model interpretability methods
Chun Sik Chan, Huanqi Kong, and Guanqing Liang · 2022
Later among the works it cites.