Fetching the paper…
Reading the bibliography…
As attribution-based explanation methods are increasingly used to establish model trustworthiness in high-stakes situations, it is critical to ensure that these explanations are stable, e.g., robust to infinitesimal perturbations to an input.
Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Simonyan Karen, Vedaldi Andrea, Zisserman Andrew · 2014
Earlier work this paper cites.
How we analyzed the COMPAS recidivism algorithm
Mattu S, Kirchner L, Angwin J · 2016
Earlier work this paper cites.
” Why should i trust you?” Explaining the predictions of any classifier
Ribeiro Marco Tulio, Singh Sameer, Guestrin Carlos · 2016
Earlier work this paper cites.
UCI Machine Learning Repository. 2017
Dua Dheeru, Graff Casey · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg Scott M, Lee Su-In · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar Avanti, Greenside Peyton, Kundaje Anshul · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Smilkov Daniel, Thorat Nikhil, Kim Been, Viégas Fernanda B., Wattenberg Martin · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo Julius, Gilmer Justin, Muelly Michael, Goodfellow Ian, Hardt Moritz, Kim Been · 2018
Cited alongside, same era.
On the robustness of interpretability methods
Alvarez-Melis David, Jaakkola Tommi S · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Ribeiro Marco Tulio, Singh Sameer, Guestrin Carlos · 2018
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Dombrowski Ann-Kathrin, Alber Maximilian, Anders Christopher J, Ackermann Marcel, Müller Klaus-Robert, Kessel Pan · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Ghorbani Amirata, Abid Abubakar, Zou James · 2019
Cited alongside, same era.
Sam: The sensitivity of attribution methods to hyperparameters
Bansal Naman, Agarwal Chirag, Nguyen Anh · 2020
Cited alongside, same era.
Towards a unified framework for fair and stable graph representation learning
Agarwal Chirag, Lakkaraju Himabindu, Zitnik Marinka · 2021
Later among the works it cites.
From Human Explanation to Model Interpretability: A Framework Based on Weight of Evidence
Alvarez-Melis David, Kaur Harmanpreet, II Hal Daum ée, Wallach Hanna, Vaughan Jennifer Wortman · 2021
Later among the works it cites.
Regularisation of neural networks by enforcing lipschitz continuity
Gouk Henry, Frank Eibe, Pfahringer Bernhard, Cree Michael J · 2021
Later among the works it cites.
Reliable post hoc explanations: Modeling uncertainty in explainability
Slack Dylan, Hilgard Anna, Singh Sameer, Lakkaraju Himabindu · 2021
Later among the works it cites.
Axiomatic attribution for deep networks
Sundararajan Mukund, Taly Ankur, Yan Qiqi · 2021
Later among the works it cites.
Probing GNN Explainers: A Rigorous Theoretical and Empirical Analysis of GNN Explanation Methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How can we fool LIME and SHAP? Adversarial Attacks on Post hoc Explanation Methods
Slack Dylan, Hilgard Sophie, Jia Emily, Singh Sameer, Lakkaraju Himabindu · 2020
Cited alongside, same era.
Agarwal Chirag, Zitnik Marinka, Lakkaraju Himabindu · 2022
Closest in time.