Fetching the paper…
Reading the bibliography…
We propose a novel method to explain trained deep neural networks (DNNs), by distilling them into surrogate models using unsupervised clustering.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Thinking, fast and slow
Kahneman, D · 2011
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Unsupervised deep embedding for clustering analysis
Xie, J., Girshick, R., and Farhadi, A · 2015
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability
Kim, B., Khanna, R., and Koyejo, O. O · 2016
Earlier work this paper cites.
”why should I trust you?”: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Distilling a neural network into a soft decision tree
Frosst, N. and Hinton, G · 2017
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., and Sayres, R · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Cited alongside, same era.
Li, O., Liu, H., Chen, C., and Rudin, C · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Yeh, C.-K., Kim, J. S., Yen, I. E. H., and Ravikumar, P · 2018
Later among the works it cites.
Interpretable basis decomposition for visual explanation
Zhou, B., Sun, Y., Bau, D., and Torralba, A · 2018
Later among the works it cites.
Arik, S. Ö. and Pfister, T · 2019
Later among the works it cites.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Arrieta, A. B., Díaz-Rodríguez, N., Ser, J. D., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., and Herrera, F · 2019
Later among the works it cites.
Leveraging large-scale uncurated data for unsupervised pre-training of visual features
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M · 2018
Cited alongside, same era.
Please Stop Explaining Black Box Models for High Stakes Decisions
Rudin, C · 2018
Cited alongside, same era.
Learning global additive explanations for neural nets using model distillation
Tan, S., Caruana, R., Hooker, G., Koch, P., and Gordo, A · 2018
Cited alongside, same era.
Caron, M., Bojanowski, P., Mairal, J., and Joulin, A · 2019
Later among the works it cites.
Towards automatic concept-based explanations
Ghorbani, A., Wexler, J., Zou, J., and Kim, B · 2019
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
Yeh, C.-K., Kim, B., Arik, S. O., Li, C.-L., Pfister, T., and Ravikumar, P · 2019
Later among the works it cites.
Rl-lim: Reinforcement learning-based locally interpretable modeling
Yoon, J., Arik, S. O., and Pfister, T · 2019
Later among the works it cites.