Fetching the paper…
Reading the bibliography…
Recently, interpretable models called self-explaining models (SEMs) have been proposed with the goal of providing interpretability robustness.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Binder, A., Montavon, G., Lapuschkin, S., Müller, K.-R., and Samek, W · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
This looks like that: deep learning for interpretable image recognition
Chen, C., Li, O., Barnett, A., Su, J., and Rudin, C · 2018
Cited alongside, same era.
Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions
Li, O., Liu, H., Chen, C., and Rudin, C · 2018
Cited alongside, same era.
On the robustness of interpretability methods
Alvarez-Melis, D. and Jaakkola, T. S
Cited in the paper.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D. A. and Jaakkola, T
Cited in the paper.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Later among the works it cites.
Interpretation of neural networks is fragile
Ghorbani, A., Abid, A., and Zou, J · 2019
Closest in time.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…