Fetching the paper…
Reading the bibliography…
We argue that robustness of explanations---i.e., that similar inputs should give rise to similar explanations---is a key desideratum for interpretability.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Practical Bayesian Optimization of Machine Learning Algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan Prescott · 2012
Earlier work this paper cites.
{UCI} Machine Learning Repository, 2013
Lichman, Moshe and Bache, Kevin · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, Karen, Vedaldi, Andrea, and Zisserman, Andrew · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, Matthew D and Fergus, Rob · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, Sebastian, Binder, Alexander, Montavon, Grégoire, Klauschen, Frederick, Müller, Klaus Robert, and Samek, Wojciech · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Earlier work this paper cites.
”Why Should I Trust You?”: Explaining the Predictions of Any Classifier
Ribeiro, Marco Tulio, Singh, Sameer, and Guestrin, Carlos · 2016
Cited alongside, same era.
Not just a black box: Learning important features through propagating activation differences
Shrikumar, Avanti, Greenside, Peyton, Shcherbina, Anna, and Kundaje, Anshul · 2016
Cited alongside, same era.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
Alvarez-Melis, David and Jaakkola, Tommi S · 2017
Cited alongside, same era.
Formal Guarantees on the Robustness of a Classifier against Adversarial Manipulation
Hein, Matthias and Andriushchenko, Maksym · 2017
Cited alongside, same era.
The (Un)reliability of saliency methods
Kindermans, P.-J., Hooker, S, Adebayo, J, Alber, M, Schütt, K.˜T., Dähne, S, Erhan, D, and Kim, B · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, Scott and Lee, Su-In · 2017
Later among the works it cites.
Selvaraju, Ramprasaath R., Das, Abhishek, Vedantam, Ramakrishna, Cogswell, Michael, Parikh, Devi, and Batra, Dhruv · 2017
Later among the works it cites.
Axiomatic attribution for deep networks
Sundararajan, Mukund, Taly, Ankur, and Yan, Qiqi · 2017
Later among the works it cites.
Towards Robust Interpretability with Self-explaining Neural Networks
Alvarez-Melis, David and Jaakkola, Tommi S · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Provable defenses against adversarial examples via the convex outer adversarial polytope
Kolter, J Zico and Wong, Eric · 2017
Cited alongside, same era.
Raghunathan, Aditi, Steinhardt, Jacob, and Liang, Percy · 2018
Closest in time.
Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
Weng, Tsui-Wei, Zhang, Huan, Chen, Pin-Yu, Yi, Jinfeng, Su, Dong, Gao, Yupeng, Hsieh, Cho-Jui, and Daniel, Luca · 2018
Closest in time.