Fetching the paper…
Reading the bibliography…
We consider objective evaluation measures of saliency explanations for complex black-box machine learning models.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and MÞller, K.-R · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Explaining prediction models and individual predictions with feature contributions
Štrumbelj, E. and Kononenko, I · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Principles of explanatory debugging to personalize interactive machine learning
Kulesza, T., Burnett, M., Wong, W.-K., and Stumpf, S · 2015
Earlier work this paper cites.
Algorithmic transparency via quantitative input influence: Theory and experiments with learning systems
Datta, A., Sen, S., and , Y. Z · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Müller, K.-R · 2016
Earlier work this paper cites.
Not just a black box: Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., Shcherbina, A., and Kundaje, A · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Real time image saliency for black box classifiers
Dabkowski, P. and Gal, Y · 2017
Earlier work this paper cites.
A roadmap for a rigorous science of interpretability
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Miller, T · 2017
Cited alongside, same era.
Methods for interpreting and understanding deep neural networks
Montavon, G., Samek, W., and Müller, K.-R · 2017
Cited alongside, same era.
Ross, A. S. and Doshi-Velez, F · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantamand, R., Parikh, D., and Parikh, D · 2017
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Later among the works it cites.
Patternnet and patternlrp–improving the interpretability of neural networks
Kindermans, P.-J., Schütt, K. T., Alber, M., Müller, K.-R., and Dähne, S · 2018
Later among the works it cites.
Towards robust neural networks via random self-ensemble
Liu, X., Cheng, M., Zhang, H., and Hsieh, C.-J · 2018
Later among the works it cites.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk, V., Das, A., and Saenko, K · 2018
Later among the works it cites.
Model agnostic supervised local explanations
Plumb, G., Molitor, D., and Talwalkar, A. S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Cited alongside, same era.
Certifiable distributional robustness with principled adversarial training
Sinha, A., Namkoong, H., and Duchi, J · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Visualizing deep neural network decisions: Prediction difference analysis
Zintgraf, L. M., Cohen, T. S., Adel, T., and Welling, M · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Cited alongside, same era.
On the robustness of interpretability methods
Alvarez-Melis, D. and Jaakkola, T. S · 2018
Cited alongside, same era.
Raghunathan, A., Steinhardt, J., and Liang, P · 2018
Later among the works it cites.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Wong, E. and Kolter, Z · 2018
Later among the works it cites.
Representer point selection for explaining deep neural networks
Yeh, C., Kim, J. S., Yen, I. E., and Ravikumar, P · 2018
Later among the works it cites.
Interpretable deep learning under fire
Zhang, X., Wang, N., Ji, S., Shen, H., and Wang, T · 2018
Later among the works it cites.
Explaining image classifiers by counterfactual generation
Chang, C.-H., Creager, E., Goldenberg, A., and Duvenaud, D · 2019
Closest in time.
Certified adversarial robustness via randomized smoothing
Cohen, J. M., Rosenfeld, E., and Kolter, J. Z · 2019
Closest in time.
Counterfactual visual explanations
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., and Lee, S · 2019
Closest in time.
The (un) reliability of saliency methods
Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., and Kim, B · 2019
Closest in time.
Towards robust, locally linear deep networks
Lee, G.-H., Alvarez-Melis, D., and Jaakkola, T. S · 2019
Closest in time.
Gnn explainer: A tool for post-hoc explanation of graph neural networks
Ying, R., Bourgeois, D., You, J., Zitnik, M., and Leskovec, J · 2019
Closest in time.