Fetching the paper…
Reading the bibliography…
With machine learning models being used for more sensitive applications, we rely on interpretability methods to prove that no discriminating attributes were used for classification.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R · 1992
Earlier work this paper cites.
Consensus inference in neuroimaging
Hansen, L. K., Nielsen, F. Å., Strother, S. C., and Lange, N · 2001
Earlier work this paper cites.
Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A · 2016
Earlier work this paper cites.
The (un)reliability of saliency methods
Kindermans, P.-J., Hooker, S., Adebayo, J., Brain, G., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., and Kim, B · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep taylor decomposition
Montavon, G., Lapuschkin, S., Binder, A., Samek, W., and Müller, K.-R · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Defense against adversarial attacks using high-level representation guided denoiser
Liao, F., Liang, M., Dong, Y., Pang, T., Hu, X., and Zhu, J · 2018
Later among the works it cites.
A human-grounded evaluation benchmark for local explanations of machine learning
Mohseni, S. and Ragan, E. D · 2018
Later among the works it cites.
Structuring Neural Networks for More Explainable Predictions
Rieger, L., Chormai, P., Montavon, G., Hansen, L. K., and Müller, K.-R · 2018
Later among the works it cites.
Fairwashing: the risk of rationalization
Aïvodji, U., Arai, H., Fortineau, O., Gambs, S., Hara, S., and Tapp, A · 2019
Later among the works it cites.
Explanations can be manipulated and geometry is to blame
Dombrowski, A.-K., Alber, M., Anders, C. J., Ackermann, M., Müller, K.-R., and Kessel, P · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tramèr, F., Kurakin, A., Papernot, N., Goodfellow, I., Boneh, D., and McDaniel, P · 2017
Cited alongside, same era.
Visualizing deep neural network decisions: Prediction difference analysis
Zintgraf, L. M., Cohen, T. S., Adel, T., and Welling, M · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Ancona, M., Ceolini, E., Oztireli, C., and Gross, M · 2018
Cited alongside, same era.
Explaining image classifiers by counterfactual generation
Chang, C.-H., Creager, E., Goldenberg, A., and Duvenaud, D · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Cited alongside, same era.
Interpretable convolutional neural networks
Zhang, Q., Nian Wu, Y., and Zhu, S.-C
Cited in the paper.
Later among the works it cites.
Interpretation of neural networks is fragile
Ghorbani, A., Abid, A., and Zou, J · 2019
Later among the works it cites.
Interpretability in intelligent systems–a new concept?
Hansen, L. K. and Rieger, L · 2019
Later among the works it cites.
Fooling neural network interpretations via adversarial model manipulation
Heo, J., Joo, S., and Moon, T · 2019
Later among the works it cites.
Improving adversarial robustness via promoting ensemble diversity
Pang, T., Xu, K., Du, C., Chen, N., and Zhu, J · 2019
Later among the works it cites.
On the (in) fidelity and sensitivity of explanations
Yeh, C.-K., Hsieh, C.-Y., Suggala, A., Inouye, D. I., and Ravikumar, P. K · 2019
Later among the works it cites.
Evaluating and aggregating feature-based model explanations
Bhatt, U., Weller, A., and Moura, J. M · 2020
Closest in time.