Fetching the paper…
Reading the bibliography…
As AI technology is increasingly applied to high-impact, high-risk domains, there have been a number of new methods aimed at making AI models more human interpretable.
Herlocker, J.L., Konstan, J.A., Riedl, J.: Explaining collaborative filtering recommendations. In: CSCW (2000)
2000
Earlier work this paper cites.
Dzindolet, M.T., Peterson, S.A., Pomranky, R.A., Pierce, L.G., Beck, H.P.: The role of trust in automation reliance. In: IJHCS (2003)
2003
Earlier work this paper cites.
Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The Caltech-UCSD Birds-200-2011 dataset. Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)
2011
Earlier work this paper cites.
Simonyan, K., Vedaldi, A., Zisserman, A.: Deep inside convolutional networks: Visualising image classification models and saliency maps. In: ICLR Workshops (2014)
2014
Earlier work this paper cites.
Zeiler, M.D., Fergus, R.: Visualizing and understanding convolutional networks. In: ECCV (2014)
2014
Earlier work this paper cites.
Zhang, P., Wang, J., Farhadi, A., Hebert, M., Parikh, D.: Predicting failures of vision systems. In: CVPR (2014)
2014
Earlier work this paper cites.
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.R., Samek, W.: On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PloS one (2015)
2015
Earlier work this paper cites.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet Large Scale Visual Recognition Challenge. IJCV (2015)
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
2016
Earlier work this paper cites.
Lakkaraju, H., Bach, S.H., Leskovec, J.: Interpretable decision sets: A joint framework for description and prediction. In: KDD (2016)
2016
Earlier work this paper cites.
Ribeiro, M.T., Singh, S., Guestrin, C.: “Why should I trust you?”: Explaining the predictions of any classifier. In: KDD (2016)
2016
Earlier work this paper cites.
Zhang, J., Lin, Z., Brandt, J., Shen, X., Sclaroff, S.: Top-down neural attention by excitation backprop. In: ECCV (2016)
2016
Earlier work this paper cites.
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: Learning deep features for discriminative localization. In: CVPR (2016)
2016
Earlier work this paper cites.
Bau, D., Zhou, B., Khosla, A., Oliva, A., Torralba, A.: Network dissection: Quantifying interpretability of deep visual representations. In: CVPR (2017)
2017
Earlier work this paper cites.
Fong, R., Vedaldi, A.: Interpretable explanations of black boxes by meaningful perturbation. In: ICCV (2017)
2017
Earlier work this paper cites.
Koh, P.W., Liang, P.: Understanding black-box predictions via influence functions. In: ICML (2017)
2017
Earlier work this paper cites.
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad-CAM: Visual explanations from deep networks via gradient-based localization. In: ICCV (2017)
2017
Earlier work this paper cites.
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., Kim, B.: Sanity checks for saliency maps. In: NeurIPS (2018)
2018
Earlier work this paper cites.
Fong, R., Vedaldi, A.: Net2Vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks. In: CVPR (2018)
2018
Earlier work this paper cites.
Gilpin, L.H., Bau, D., Yuan, B.Z., Bajwa, A., Specter, M., Kagal, L.: Explaining explanations: An overview of interpretability of machine learning. In: DSAA (2018)
2018
Earlier work this paper cites.
Lage, I., Ross, A.S., Kim, B., Gershman, S.J., Doshi-Velez, F.: Human-in-the-loop interpretability prior. In: NeurIPS (2018)
2018
Earlier work this paper cites.
Lipton, Z.C.: The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue (2018)
2018
Earlier work this paper cites.
Petsiuk, V., Das, A., Saenko, K.: RISE: Randomized input sampling for explanation of black-box models. In: BMVC (2018)
2018
Earlier work this paper cites.
Wang, P., Vasconcelos, N.: Towards realistic predictors. In: ECCV (2018)
2018
Earlier work this paper cites.
Yeh, C.K., Kim, J., Yen, I.E.H., Ravikumar, P.K.: Representer point selection for explaining deep neural networks. In: NeurIPS (2018)
2018
Earlier work this paper cites.
Zhou, B., Sun, Y., Bau, D., Torralba, A.: Interpretable basis decomposition for visual explanation. In: ECCV (2018)
2018
Earlier work this paper cites.
Bau, D., Zhu, J.Y., Wulff, J., Peebles, W., Strobelt, H., Zhou, B., Torralba, A.: Seeing what a GAN cannot generate. In: ICCV (2019)
2019
Earlier work this paper cites.
Biessmann, F., Refiano, D.I.: A psychophysics approach for quantitative comparison of interpretable computer vision models (2019)
2019
Cited alongside, same era.
Brendel, W., Bethge, M.: Approximating CNNs with bag-of-local-features models works surprisingly well on imagenet. In: ICLR (2019)
2019
Cited alongside, same era.
Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., Su, J.K.: This looks like that: Deep learning for interpretable image recognition. In: NeurIPS (2019)
2019
Cited alongside, same era.
Fong, R., Patrick, M., Vedaldi, A.: Understanding deep networks via extremal perturbations and smooth masks. In: ICCV (2019)
2019
Cited alongside, same era.
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., Lee, S.: Counterfactual visual explanations. In: ICML (2019)
2019
Cited alongside, same era.
Borowski, J., Zimmermann, R.S., Schepers, J., Geirhos, R., Wallis, T.S.A., Bethge, M., Brendel, W.: Exemplary natural images explain CNN activations better than state-of-the-art feature visualization. In: ICLR (2021)
2021
Closest in time.
Böhle, M., Fritz, M., Schiele, B.: Convolutional dynamic alignment networks for interpretable classifications. In: CVPR (2021)
2021
Closest in time.
Chen, V., Li, J., Kim, J.S., Plumb, G., Talwalkar, A.: Towards connecting use cases and methods in interpretable machine learning. In: ICML Workshop on Human Interpretability in Machine Learning (2021)
2021
Closest in time.
Ehsan, U., Wintersberger, P., Liao, Q.V., Mara, M., Streit, M., Wachter, S., Riener, A., Riedl, M.O.: Operationalizing human-centered perspectives in explainable AI. CHI Extended Abstracts (2021)
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gunning, D., Aha, D.: DARPA’s explainable artificial intelligence (XAI) program. AI Magazine (2019)
2019
Cited alongside, same era.
Hooker, S., Erhan, D., Kindermans, P.J., Kim, B.: A benchmark for interpretability methods in deep neural networks. In: NeurIPS (2019)
2019
Cited alongside, same era.
Kunkel, J., Donkers, T., Michael, L., Barbu, C.M., Ziegler, J.: Let me explain: Impact of personal and impersonal explanations on trust in recommender systems. In: CHI (2019)
2019
Cited alongside, same era.
Lage, I., Chen, E., He, J., Narayanan, M., Kim, B., Gershman, S.J., Doshi-Velez, F.: Human evaluation of models built for interpretability. In: HCOMP (2019)
2019
Cited alongside, same era.
Lai, V., Tan, C.: On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In: FAccT (2019)
2019
Cited alongside, same era.
Samek, W., Montavon, G., Vedaldi, A., Hansen, L.K., Müller, K. (eds.): Explainable AI: Interpreting, Explaining and Visualizing Deep Learning. Springer (2019)
2019
Cited alongside, same era.
Schaffer, J., O’Donovan, J., Michaelis, J., Raglin, A., Höllerer, T.: I can do better than your AI: Expertise and explanations. In: IUI (2019)
2019
Cited alongside, same era.
Fel, T., Colin, J., Cadène, R., Serre, T.: What I cannot predict, I do not understand: A human-centered evaluation framework for explainability methods (2021)
2021
Closest in time.
Hoffmann, A., Fanconi, C., Rade, R., Kohler, J.: This looks like that… does it? Shortcomings of latent space prototype interpretability in deep networks. In: ICML Workshop on Theoretic Foundation, Criticism, and Application Trend of Explainable AI (2021)
2021
Closest in time.
Kim, B., Reif, E., Wattenberg, M., Bengio, S., Mozer, M.C.: Neural networks trained on natural scenes exhibit gestalt closure. Computational Brain & Behavior (2021)
2021
Closest in time.
Liao, Q.V., Varshney, K.R.: Human-centered explainable AI (XAI): From algorithms to user experiences. arXiv (2021)
2021
Closest in time.
Margeloiu, A., Ashman, M., Bhatt, U., Chen, Y., Jamnik, M., Weller, A.: Do concept bottleneck models learn as intended? In: ICLR Workshop on Responsible AI (2021)
2021
Closest in time.
Nauta, M., van Bree, R., Seifert, C.: Neural prototype trees for interpretable fine-grained image recognition. In: CVPR (2021)
2021
Closest in time.
Nguyen, G., Kim, D., Nguyen, A.: The effectiveness of feature attribution methods and its correlation with automatic evaluation scores. In: NeurIPS (2021)
2021
Closest in time.
Poppi, S., Cornia, M., Baraldi, L., Cucchiara, R.: Revisiting the evaluation of class activation mapping for explainability: A novel metric and experimental analysis. In: CVPR Workshop on Responsible Computer Vision (2021)
2021
Closest in time.
Poursabzi-Sangdeh, F., Goldstein, D.G., Hofman, J.M., Wortman Vaughan, J.W., Wallach, H.: Manipulating and measuring model interpretability. In: CHI (2021)
2021
Closest in time.
Rudin, C., Chen, C., Chen, Z., Huang, H., Semenova, L., Zhong, C.: Interpretable machine learning: Fundamental principles and 10 grand challenges. In: Statistics Surveys (2021)
2021
Closest in time.
Shitole, V., Li, F., Kahng, M., Tadepalli, P., Fern, A.: One explanation is not enough: Structured attention graphs for image classification. In: NeurIPS (2021)
2021
Closest in time.
Zimmermann, R.S., Borowski, J., Geirhos, R., Bethge, M., Wallis, T.S.A., Brendel, W.: How well do feature visualizations support causal understanding of CNN activations? In: NeurIPS (2021)
2021
Closest in time.
Agarwal, C., D’souza, D., Hooker, S.: Estimating example difficulty using variance of gradients. In: CVPR (2022)
2022
Closest in time.
Bylinskii, Z., Herman, L., Hertzmann, A., Hutka, S., Zhang, Y.: Towards better user studies in computer graphics and vision. arXiv (2022)
2022
Closest in time.
Böhle, M., Fritz, M., Schiele, B.: B-cos networks: Alignment is all we need for interpretability. In: CVPR (2022)
2022
Closest in time.
Donnelly, J., Barnett, A.J., Chen, C.: Deformable ProtoPNet: An interpretable image classifier using deformable prototypes. In: CVPR (2022)
2022
Closest in time.
Dubey, A., Radenovic, F., Mahajan, D.: Scalable interpretability via polynomials. arXiv (2022)
2022
Closest in time.
Ehsan, U., Wintersberger, P., Liao, Q.V., Watkins, E.A., Manger, C., III, H.D., Riener, A., Riedl, M.O.: Human-Centered Explainable AI (HCXAI): Beyond opening the black-box of AI. CHI Extended Abstracts (2022)
2022
Closest in time.
Radenovic, F., Dubey, A., Mahajan, D.: Neural basis models for interpretability. arXiv (2022)
2022
Closest in time.
Ramaswamy, V.V., Kim, S.S.Y., Fong, R., Russakovsky, O.: Overlooked factors in concept-based explanations: Dataset choice, concept salience, and human capability. arXiv (2022)
2022
Closest in time.
Ramaswamy, V.V., Kim, S.S.Y., Meister, N., Fong, R., Russakovsky, O.: ELUDE: Generating interpretable explanations via a decomposition into labelled and unlabelled features. arXiv (2022)
2022
Closest in time.
Vandenhende, S., Mahajan, D., Radenovic, F., Ghadiyaram, D.: Making heads or tails: Towards semantically consistent visual counterfactuals. In: ECCV (2022)
2022
Closest in time.