Fetching the paper…
Reading the bibliography…
Interpretability is an important area of research for safe deployment of machine learning systems.
Fooling neural network interpretations via adversarial model manipulation
Heo, J., Joo, S., and Moon, T. (2019) · 1902
Earlier work this paper cites.
Learning perceptually-aligned representations via adversarial robustness
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Tran, B., and Madry, A. (2019) · 1906
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A. C., and Vincent, P. (2009) · 2009
Earlier work this paper cites.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and Müller, K.-R. (2010) · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A. (2013) · 2013
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Bourdev, L. D., Girshick, R. B., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R. (2014) · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Bernstein, M. S., Fei-Fei, L., Berg, A. C., and Khosla, A. (2015) · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
Lakkaraju, H., Bach, S. H., and Leskovec, J. (2016) · 2016
Cited alongside, same era.
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016) · 2016
Cited alongside, same era.
Grad-cam: Why did you say that?
Selvaraju, R. R., Das, A., Vedantam, R., Cogswell, M., Parikh, D., and Batra, D. (2016) · 2016
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B. (2017) · 2017
Cited alongside, same era.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. and Vedaldi, A. (2017) · 2017
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D. and Jaakkola, T. S. (2018) · 2018
Later among the works it cites.
Towards better understanding of gradient-based attribution methods for deep neural networks
Ancona, M. B., Ceolini, E., Oztireli, C., and Gross, M. H. (2018) · 2018
Later among the works it cites.
Interpretation of neural networks is fragile
Ghorbani, A., Abid, A., and Zou, J. Y. (2018) · 2018
Later among the works it cites.
Evaluating feature importance estimates
Hooker, S., Erhan, D., Kindermans, P., and Kim, B. (2018) · 2018
Later among the works it cites.
Excessive invariance causes adversarial vulnerability
Jacobsen, J.-H., Behrmann, J., Zemel, R., and Bethge, M. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating the visualization of what a deep neural network has learned
Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Müller, K.-R. (2017) · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F. B., and Wattenberg, M. (2017) · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q. (2017) · 2017
Cited alongside, same era.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A. (2017) · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I. J., Hardt, M., and Kim, B. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C. J., Wexler, J., Viégas, F. B., and Sayres, R. (2018) · 2018
Later among the works it cites.
Human-in-the-loop interpretability prior
Lage, I., Ross, A., Gershman, S. J., Kim, B., and Doshi-Velez, F. (2018) · 2018
Later among the works it cites.
Narayanan, M., Chen, E., He, J., Kim, B., Gershman, S., and Doshi-Velez, F. (2018) · 2018
Later among the works it cites.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Nie, W., Zhang, Y., and Patel, A. (2018) · 2018
Later among the works it cites.
Manipulating and measuring model interpretability
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Vaughan, J. W., and Wallach, H. (2018) · 2018
Later among the works it cites.