Fetching the paper…
Reading the bibliography…
A critical problem in the field of post hoc explainability is the lack of a common foundational goal among methods.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew Zeiler and Robert Fergus · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2015
Earlier work this paper cites.
Learning deconvolution network for semantic segmentation
Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han · 2015
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
"Why should I trust you?" Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
SmoothGrad: Removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Grad-CAM: Visual explanations from deep networks via gradient-based localization
Ramprasaath Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Artificial intelligence in healthcare
Kun-Hsing Yu, Andrew L Beam, and Isaac S Kohane · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Cited alongside, same era.
Knowledge transfer with Jacobian matching
Suraj Srinivas and François Fleuret · 2018
Cited alongside, same era.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi Jaakkola · 2018
Cited alongside, same era.
Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks
Full-gradient representation for neural network visualization
Suraj Srinivas and François Fleuret · 2019
Later among the works it cites.
Home equity line of credit (HELOC) dataset
FICO · 2019
Later among the works it cites.
Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Later among the works it cites.
Captum: A unified and generic model interpretability library for PyTorch, 2020
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson · 2020
Later among the works it cites.
Artificial intelligence and law
Robert Walters and Marko Novak · 2021
Later among the works it cites.
Towards the unification and robustness of perturbation and gradient based explanations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth Balasubramanian · 2018
Cited alongside, same era.
Life expectancy dataset
World Health Organization (WHO) · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary Lipton · 2018
Cited alongside, same era.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Weili Nie, Yang Zhang, and Ankit Patel · 2018
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximilian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel · 2019
Cited alongside, same era.
Sushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay, Steven Wu, and Himabindu Lakkaraju · 2021
Later among the works it cites.
Explaining by removing: A unified framework for model explanation
Ian Covert, Scott Lundberg, and Su-In Lee · 2021
Later among the works it cites.
Rethinking the role of gradient-based attribution methods for model interpretability
Suraj Srinivas and Francois Fleuret · 2021
Later among the works it cites.
AI in finance: Challenges, techniques, and opportunities
Longbing Cao · 2022
Closest in time.
The disagreement problem in explainable machine learning: A practitioner’s perspective
Satyapriya Krishna*, Tessa Han*, Alex Gu, Javin Pombra, Shahin Jabbari, Steven Wu, and Himabindu Lakkaraju · 2022
Closest in time.
Fairness via explanation quality: Evaluating disparities in the quality of post hoc explanations
Jessica Dai, Sohini Upadhyay, Ulrich Aivodji, Stephen Bach, and Himabindu Lakkaraju · 2022
Closest in time.