Fetching the paper…
Reading the bibliography…
The rising popularity of explainable artificial intelligence (XAI) to understand high-performing black boxes raised the question of how to evaluate explanations of machine learning (ML) models.
The measurement of observer agreement for categorical data
Landis, J. R., and Koch, G. G · 1977
Earlier work this paper cites.
Explanation in Second Generation Expert Systems
Swartout, W. R., and Moore, J. D · 1993
Earlier work this paper cites.
Survey and critique of techniques for extracting rules from trained artificial neural networks
Andrews, R., Diederich, J., and Tickle, A. B · 1995
Earlier work this paper cites.
Extracting tree-structured representations of trained networks
Craven, M. W., and Shavlik, J. W · 1995
Earlier work this paper cites.
Abductive inference: Computation, philosophy, technology
Josephson, J. R., and Josephson, S. G · 1996
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Friedman, J. H · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
An explainable artificial intelligence system for small-unit tactical behavior
Van Lent, M., Fisher, W., and Mancuso, M · 2004
Earlier work this paper cites.
Usability evaluation considered harmful (some of the time)
Greenberg, S., and Buxton, B · 2008
Earlier work this paper cites.
Visualizing data using t-sne
van der Maaten, L., and Hinton, G · 2008
Earlier work this paper cites.
An empirical evaluation of the comprehensibility of decision table, tree and rule based predictive models
Huysmans, J., Dejaeger, K., Mues, C., Vanthienen, J., and Baesens, B · 2011
Earlier work this paper cites.
Too much, too little, or just right? ways explanations impact end users’ mental models
Kulesza, T., Stumpf, S., Burnett, M., Yang, S., Kwan, I., and Wong, W.-K · 2013
Earlier work this paper cites.
Interpretable semantic vectors from a joint model of brain- and text- based meaning
Fyshe, A., Talukdar, P. P., Murphy, B., and Mitchell, T. M · 2014
Earlier work this paper cites.
Utilizing temporal patterns for estimating uncertainty in interpretable early decision making
Ghalwash, M. F., Radosavljevic, V., and Obradovic, Z · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation
Goldstein, A., Kapelner, A., Bleich, J., and Pitkin, E · 2015
Earlier work this paper cites.
Scalable and interpretable data representation for high-dimensional, complex data
Kim, B., Patel, K., Rostamizadeh, A., and Shah, J. A · 2015
Earlier work this paper cites.
Principles of explanatory debugging to personalize interactive machine learning
Kulesza, T., Burnett, M., Wong, W.-K., and Stumpf, S · 2015
Earlier work this paper cites.
Explaining recommendations: Design and evaluation
Tintarev, N., and Masthoff, J · 2015
Earlier work this paper cites.
The human kernel
Wilson, A. G., Dann, C., Lucas, C., and Xing, E. P · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability
Kim, B., Koyejo, O., and Khanna, R · 2016
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
Lakkaraju, H., Bach, S. H., and Leskovec, J · 2016
Earlier work this paper cites.
Confusions over time: An interpretable bayesian model to characterize trends in decision making
Lakkaraju, H., and Leskovec, J · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Interpretable nonlinear dynamic modeling of neural trajectories
Zhao, Y., and Park, I. M · 2016
Earlier work this paper cites.
A unified view of gradient-based attribution methods for Deep Neural Networks
Ancona, M., Öztireli, C., Ceolini, E., and Gross, M · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Explanation and justification in machine learning: A survey
Biran, O., and Cotton, C · 2017
Earlier work this paper cites.
Improving interpretability of deep neural networks with semantic information
Dong, Y., Su, H., Zhu, J., and Zhang, B · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. C., and Vedaldi, A · 2017
Earlier work this paper cites.
The Promise and Peril of Human Evaluation for Model Interpretability
Herman, B · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled and interpretable representations from sequential data
Hsu, W., Zhang, Y., and Glass, J. R · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M., and Lee, S · 2017
Earlier work this paper cites.
Explainable ai: Beware of inmates running the asylum or: How i learnt to stop worrying and love the social and behavioural sciences
Miller, T., Howe, P., and Sonenberg, L · 2017
Earlier work this paper cites.
A systematic review and taxonomy of explanations in decision support and recommender systems
Nunes, I., and Jannach, D · 2017
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
SVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations
Ross, A. S., Hughes, M. C., and Doshi-Velez, F · 2017
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Müller, K.-R · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
Interpretable predictions of tree-based ensembles via actionable feature tweaking
Tolomei, G., Silvestri, F., Haines, A., and Lalmas, M · 2017
Earlier work this paper cites.
Optimized risk scores
Ustun, B., and Rudin, C · 2017
Earlier work this paper cites.
Interpretable transformations with encoder-decoder networks
Worrall, D. E., Garbin, S. J., Turmukhambetov, D., and Brostow, G. J · 2017
Earlier work this paper cites.
Towards deep interpretability (MUS-ROVER II): learning hierarchical representations of tonal music
Yu, H., and Varshney, L. R · 2017
Earlier work this paper cites.
Evaluating everyday explanations
Zemla, J. C., Sloman, S., Bechlivanidis, C., and Lagnado, D. A · 2017
Earlier work this paper cites.
Growing interpretable part graphs on convnets via multi-shot learning
Zhang, Q., Cao, R., Wu, Y. N., and Zhu, S · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (xai)
Adadi, A., and Berrada, M · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
Discovering interpretable representations for both deep generative and discriminative models
Adel, T., Ghahramani, Z., and Weller, A · 2018
Earlier work this paper cites.
Auditing black-box models for indirect influence
Adler, P., Falk, C., Friedler, S. A., Nix, T., Rybeck, G., Scheidegger, C., Smith, B., and Venkatasubramanian, S · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D., and Jaakkola, T. S · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Camburu, O., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P · 2018
Earlier work this paper cites.
Neural attentional rating regression with review-level explanations
Chen, C., Zhang, M., Liu, Y., and Ma, S · 2018
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation
Chen, J., Song, L., Wainwright, M. J., and Jordan, M. I · 2018
Earlier work this paper cites.
Exact and consistent interpretation for piecewise linear neural networks: A closed form solution
Chu, L., Hu, X., Hu, J., Wang, L., and Pei, J · 2018
Earlier work this paper cites.
Learning to act properly: Predicting and explaining affordances from images
Chuang, C., Li, J., Torralba, A., and Fidler, S · 2018
Earlier work this paper cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Dhurandhar, A., Chen, P., Luss, R., Tu, C., Ting, P., Shanmugam, K., and Das, P · 2018
Earlier work this paper cites.
Considerations for Evaluation and Generalization in Interpretable Machine Learning
Doshi-Velez, F., and Kim, B · 2018
Earlier work this paper cites.
Towards explanation of dnn-based prediction with guided feature inversion
Du, M., Liu, N., Song, Q., and Hu, X · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Fong, R., and Vedaldi, A · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., and Kagal, L · 2018
Earlier work this paper cites.
A survey of methods for explaining black box models
Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., and Pedreschi, D · 2018
Earlier work this paper cites.
Explaining deep learning models - A bayesian non-parametric approach
Guo, W., Huang, S., Tao, Y., Xing, X., and Lin, L · 2018
Earlier work this paper cites.
A data-driven analysis of workers’ earnings on amazon mechanical turk
Hara, K., Adams, A., Milland, K., Savage, S., Callison-Burch, C., and Bigham, J. P · 2018
Earlier work this paper cites.
Uncertainty-aware attention for reliable interpretation and prediction
Heo, J., Lee, H., Kim, S., Lee, J., Kim, K. J., Yang, E., and Hwang, S. J · 2018
Earlier work this paper cites.
Honegger, M · 2018
Earlier work this paper cites.
Interpretable word embeddings for medical domain
Jha, K., Wang, Y., Xun, G., and Zhang, A · 2018
Earlier work this paper cites.
Explainable time series tweaking via irreversible and reversible temporal transformations
Karlsson, I., Rebane, J., Papapetrou, P., and Gionis, A · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C. J., Wexler, J., Viégas, F. B., and Sayres, R · 2018
Earlier work this paper cites.
Disentangling by factorising
Kim, H., and Mnih, A · 2018
Earlier work this paper cites.
Learning how to explain neural networks: Patternnet and patternattribution
Kindermans, P., Schütt, K. T., Alber, M., Müller, K., Erhan, D., Kim, B., and Dähne, S · 2018
Earlier work this paper cites.
Explaining multi-criteria decision aiding models with an extended shapley value
Labreuche, C., and Fossier, S · 2018
Earlier work this paper cites.
Human-in-the-loop interpretability prior
Lage, I., Ross, A. S., Gershman, S. J., Kim, B., and Doshi-Velez, F · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Lipton, Z. C · 2018
Earlier work this paper cites.
On interpretation of network embedding via taxonomy induction
Liu, N., Huang, X., Li, J., and Hu, X · 2018
Earlier work this paper cites.
Contextual outlier interpretation
Liu, N., Shin, D., and Hu, X · 2018
Earlier work this paper cites.
Adversarial detection with model interpretation
Liu, N., Yang, H., and Hu, X · 2018
Earlier work this paper cites.
Beyond polarity: Interpretable financial sentiment analysis with hierarchical query-driven attention
Luo, L., Ao, X., Pan, F., Wang, J., Zhao, T., Yu, N., and He, Q · 2018
Earlier work this paper cites.
Transparency by design: Closing the gap between performance and interpretability in visual reasoning
Mascharka, D., Tran, P., Soklaski, R., and Majumdar, A · 2018
Earlier work this paper cites.
Methods for interpreting and understanding deep neural networks
Montavon, G., Samek, W., and Müller, K.-R · 2018
Earlier work this paper cites.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Nie, W., Zhang, Y., and Patel, A · 2018
Earlier work this paper cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Park, D. H., Hendricks, L. A., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., and Rohrbach, M · 2018
Earlier work this paper cites.
Explanation mining: Post hoc interpretability of latent factor models for recommendation systems
Peake, G., and Wang, J · 2018
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk, V., Das, A., and Saenko, K · 2018
Earlier work this paper cites.
Model agnostic supervised local explanations
Plumb, G., Molitor, D., and Talwalkar, A. S · 2018
Earlier work this paper cites.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
Pörner, N., Schütze, H., and Roth, B · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Ribeiro, M. T., Singh, S., and Guestrin, C · 2018
Earlier work this paper cites.
Perturbation-based explanations of prediction models
Robnik-Šikonja, M., and Bohanec, M · 2018
Earlier work this paper cites.
Interpretable graph-based semi-supervised learning via flows
Rustamov, R. M., and Klosowski, J. T · 2018
Earlier work this paper cites.
The internet is enabling a new kind of poorly paid hell
Semuels, A · 2018
Earlier work this paper cites.
Towards complementary explanations using deep neural networks
Silva, W., Fernandes, K., Cardoso, M. J., and Cardoso, J. S · 2018
Earlier work this paper cites.
SPINE: sparse interpretable neural embeddings
Subramanian, A., Pruthi, D., Jhamtani, H., Berg-Kirkpatrick, T., and Hovy, E. H · 2018
Earlier work this paper cites.
Neural interaction transparency (NIT): disentangling learned interactions for improved interpretability
Tsang, M., Liu, H., Purushotham, S., Murali, P., and Liu, Y · 2018
Earlier work this paper cites.
Explainable recommendation via multi-task learning in opinionated text data
Wang, N., Wang, H., Jia, Y., and Yin, Y · 2018
Earlier work this paper cites.
Multi-value rule sets for interpretable classification with feature-efficient representations
Wang, T · 2018
Earlier work this paper cites.
A reinforcement learning framework for explainable recommendation
Wang, X., Chen, Y., Yang, J., Wu, L., Wu, Z., and Xie, X · 2018
Earlier work this paper cites.
TEM: tree-enhanced embedding model for explainable recommendation
Wang, X., He, X., Feng, F., Nie, L., and Chua, T · 2018
Earlier work this paper cites.
Interpret neural networks by identifying critical data routing paths
Wang, Y., Su, H., Zhang, B., and Hu, X · 2018
Earlier work this paper cites.
Sharing deep neural network models with interpretation
Wu, H., Wang, C., Yin, J., Lu, K., and Zhu, L · 2018
Earlier work this paper cites.
Beyond sparsity: Tree regularization of deep models for interpretability
Wu, M., Hughes, M. C., Parbhoo, S., Zazzi, M., Roth, V., and Doshi-Velez, F · 2018
Earlier work this paper cites.
Towards interpretation of recommender systems with sorted explanation paths
Yang, F., Liu, N., Wang, S., and Hu, X · 2018
Earlier work this paper cites.
Interpreting CNN knowledge via an explanatory graph
Zhang, Q., Cao, R., Shi, F., Wu, Y. N., and Zhu, S · 2018
Earlier work this paper cites.
Interpretable convolutional neural networks
Zhang, Q., Wu, Y. N., and Zhu, S · 2018
Earlier work this paper cites.
Visual interpretability for deep learning: a survey
Zhang, Q.-s., and Zhu, S.-c · 2018
Earlier work this paper cites.
Interpreting neural network judgments via minimal, stable, and symbolic corrections
Zhang, X., Solar-Lezama, A., and Singh, R · 2018
Cited alongside, same era.
Unsupervised discrete sentence representation learning for interpretable neural dialog generation
Zhao, T., Lee, K., and Eskénazi, M · 2018
Cited alongside, same era.
Explaining deep neural networks with a polynomial time algorithm for shapley value approximation
Ancona, M., Öztireli, C., and Gross, M. H · 2019
Cited alongside, same era.
Explaining reinforcement learning to mere mortals: An empirical study
Anderson, A., Dodge, J., Sadarangani, A., Juozapaitis, Z., Newman, E., Irvine, J., Chattopadhyay, S., Fern, A., and Burnett, M · 2019
Cited alongside, same era.
Towards better interpretability in deep q-networks
Annasamy, R. M., and Sycara, K. P · 2019
Cited alongside, same era.
One explanation does not fit all: A toolkit and taxonomy of ai explainability techniques
Try this instead: Personalized and interpretable substitute recommendation
Chen, T., Yin, H., Ye, G., Huang, Z., Wang, Y., and Wang, M · 2020
Later among the works it cites.
Towards explainable conversational recommendation
Chen, Z., Wang, X., Xie, X., Parsana, M., Soni, A., Ao, X., and Chen, E · 2020
Later among the works it cites.
Explaining knowledge distillation by quantifying the knowledge
Cheng, X., Rao, Z., Chen, Y., and Zhang, Q · 2020
Later among the works it cites.
A Taxonomy for Human Subject Evaluation of Black-Box Explanations in XAI
Chromik, M., and Schuessler, M · 2020
Later among the works it cites.
Learning outside the black-box: The pursuit of interpretable models
Crabbé, J., Zhang, Y., Zame, W. R., and van der Schaar, M · 2020
Later among the works it cites.
Explainable data decompositions
Dalleiger, S., and Vreeken, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arya, V., Bellamy, R. K., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S. C., Houde, S., Liao, Q. V., Luss, R., Mojsilović, A., et al · 2019
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Bastings, J., Aziz, W., and Titov, I · 2019
Cited alongside, same era.
Evaluating the interpretability of the knowledge compilation map: Communicating logical statements effectively
Booth, S., Muise, C., and Shah, J · 2019
Cited alongside, same era.
Machine learning interpretability: A survey on methods and metrics
Carvalho, D. V., Pereira, E. M., and Cardoso, J. S · 2019
Cited alongside, same era.
Explaining image classifiers by counterfactual generation
Chang, C., Creager, E., Goldenberg, A., and Duvenaud, D · 2019
Cited alongside, same era.
This looks like that: Deep learning for interpretable image recognition
Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., and Su, J · 2019
Cited alongside, same era.
Scalable explanation of inferences on large graphs
Chen, C., Liu, Y., Zhang, X., and Xie, S · 2019
Cited alongside, same era.
Opportunities and challenges in explainable artificial intelligence (xai): A survey
Das, A., and Rad, P · 2020
Later among the works it cites.
A disentangling invertible interpretation network for explaining latent representations
Esser, P., Rombach, R., and Ommer, B · 2020
Later among the works it cites.
Asymmetric shapley values: incorporating causal knowledge into model-agnostic explainability
Frye, C., Rowat, C., and Feige, I · 2020
Later among the works it cites.
Fairness-aware explainable recommendation over knowledge graphs
Fu, Z., Xian, Y., Gao, R., Zhao, J., Huang, Q., Ge, Y., Xu, S., Geng, S., Shah, C., Zhang, Y., and de Melo, G · 2020
Later among the works it cites.
Simple, interpretable and stable method for detecting words with usage change across corpora
Gonen, H., Jawahar, G., Seddah, D., and Goldberg, Y · 2020
Later among the works it cites.
Interpretable deep graph generation with node-edge co-disentanglement
Guo, X., Zhao, L., Qin, Z., Wu, L., Shehu, A., and Ye, Y · 2020
Later among the works it cites.
Explaining black box predictions and unveiling data artifacts through influence functions
Han, X., Wallace, B. C., and Tsvetkov, Y · 2020
Later among the works it cites.
Interpretable and differentially private predictions
Harder, F., Bauer, M., and Park, M · 2020
Later among the works it cites.
Evaluating explainable AI: which algorithmic explanations help users predict model behavior?
Hase, P., and Bansal, M · 2020
Later among the works it cites.
Causal shapley values: Exploiting causal knowledge to explain individual predictions of complex models
Heskes, T., Sijben, E., Bucur, I. G., and Claassen, T · 2020
Later among the works it cites.
Interpretable models for understanding immersive simulations
Hoernle, N., Gal, K., Grosz, B. J., Lyons, L., Ren, A., and Rubin, A · 2020
Later among the works it cites.
Learning with interpretable structure from gated rnn
Hou, B.-J., and Zhou, Z.-H · 2020
Later among the works it cites.
Towards interpretation of pairwise learning
Huai, M., Wang, D., Miao, C., and Zhang, A · 2020
Later among the works it cites.
Interpretable and accurate fine-grained recognition via region grouping
Huang, Z., and Li, Y · 2020
Later among the works it cites.
Benchmarking deep learning interpretability in time series predictions
Ismail, A. A., Gunady, M. K., Bravo, H. C., and Feizi, S · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Jacovi, A., and Goldberg, Y · 2020
Later among the works it cites.
Self-supervised learning of interpretable keypoints from unlabelled videos
Jakab, T., Gupta, A., Bilen, H., and Vedaldi, A · 2020
Later among the works it cites.
How can I explain this to you? an empirical study of deep neural network explanation methods
Jeyakumar, J. V., Noor, J., Cheng, Y., Garcia, L., and Srivastava, M. B · 2020
Later among the works it cites.
Multi-objective molecule generation using interpretable substructures
Jin, W., Barzilay, R., and Jaakkola, T. S · 2020
Later among the works it cites.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Jin, X., Wei, Z., Du, J., Xue, X., and Ren, X · 2020
Later among the works it cites.
DACE: distribution-aware counterfactual explanation by mixed-integer linear optimization
Kanamori, K., Takagi, T., Kobayashi, K., and Arimura, H · 2020
Later among the works it cites.
A programmatic and semantic approach to explaining and debugging neural network based object detectors
Kim, E., Gopinath, D., Pasareanu, C. S., and Seshia, S. A · 2020
Later among the works it cites.
NILE : Natural language inference with faithful natural language explanations
Kumar, S., and Talukdar, P. P · 2020
Later among the works it cites.
Robust and stable black box explanations
Lakkaraju, H., Arsov, N., and Bastani, O · 2020
Later among the works it cites.
Synthesizing aspect-driven recommendation explanations from reviews
Le, T., and Lauw, H. W · 2020
Later among the works it cites.
GRACE: generating concise and informative contrastive sample to explain neural network model’s prediction
Le, T., Wang, S., and Lee, D · 2020
Later among the works it cites.
Towards falsifiable interpretability research
Leavitt, M. L., and Morcos, A · 2020
Later among the works it cites.
Evaluating explanation methods for neural machine translation
Li, J., Liu, L., Li, H., Li, G., Huang, G., and Shi, S · 2020
Later among the works it cites.
Adversarial infidelity learning for model interpretation
Liang, J., Bai, B., Cao, Y., Bai, K., and Wang, F · 2020
Later among the works it cites.
Lp-explain: Local pictorial explanation for outliers
Liu, H., Ma, F., Wang, Y., He, S., Chen, J., and Gao, J · 2020
Later among the works it cites.
Towards visually explaining variational autoencoders
Liu, W., Li, R., Zheng, M., Karanam, S., Wu, Z., Bhanu, B., Radke, R. J., and Camps, O. I · 2020
Later among the works it cites.
Parameterized explainer for graph neural network
Luo, D., Cheng, W., Xu, D., Yu, W., Zong, B., Chen, H., and Zhang, X · 2020
Later among the works it cites.
Explainable reinforcement learning through a causal lens
Madumal, P., Miller, T., Sonenberg, L., and Vetere, F · 2020
Later among the works it cites.
Explaining naive bayes and other linear classifiers with polynomial time and delay
Marques-Silva, J., Gerspacher, T., Cooper, M. C., Ignatiev, A., and Narodytska, N · 2020
Later among the works it cites.
Towards transparent and explainable attention models
Mohankumar, A. K., Nema, P., Narasimhan, S., Khapra, M. M., Srinivasan, B. V., and Ravindran, B · 2020
Later among the works it cites.
Interpretable machine learning
Molnar, C · 2020
Later among the works it cites.
Compositional explanations of neurons
Mu, J., and Andreas, J · 2020
Later among the works it cites.
Relative attributing propagation: Interpreting the comparative contributions of individual units in deep neural networks
Nam, W., Gur, S., Choi, J., Wolf, L., and Lee, S · 2020
Later among the works it cites.
Generative causal explanations of black-box classifiers
O’Shaughnessy, M. R., Canal, G., Connor, M., Rozell, C., and Davenport, M. A · 2020
Later among the works it cites.
Interpretable and personalized apprenticeship scheduling: Learning interpretable scheduling policies from heterogeneous user demonstrations
Paleja, R. R., Silva, A., Chen, L., and Gombolay, M. C · 2020
Later among the works it cites.
Explainable recommendation via interpretable feature mapping and evaluation of explainability
Pan, D., Li, X., Li, X., and Zhu, D · 2020
Later among the works it cites.
xgail: Explainable generative adversarial imitation learning for explainable human decision analysis
Pan, M., Huang, W., Li, Y., Zhou, X., and Luo, J · 2020
Later among the works it cites.
Multiresolution tensor learning for efficient and interpretable spatial analysis
Park, J. Y., Carr, K. T., Zheng, S., Yue, Y., and Yu, R · 2020
Later among the works it cites.
Explanation vs attention: A two-player game to obtain attention for VQA
Patro, B. N., Anupriy, and Namboodiri, V · 2020
Later among the works it cites.
Learning model-agnostic counterfactual explanations for tabular data
Pawelczyk, M., Broelemann, K., and Kasneci, G · 2020
Later among the works it cites.
Learning global transparent models consistent with local contrastive explanations
Pedapati, T., Balakrishnan, A., Shanmugam, K., and Dhurandhar, A · 2020
Later among the works it cites.
Regularizing black-box models for improved interpretability
Plumb, G., Al-Shedivat, M., Cabrera, Á. A., Perer, A., Xing, E. P., and Talwalkar, A · 2020
Later among the works it cites.
Explaining groups of points in low-dimensional representations
Plumb, G., Terhorst, J., Sankararaman, S., and Talwalkar, A · 2020
Later among the works it cites.
Learning to deceive with attention-based explanations
Pruthi, D., Gupta, M., Dhingra, B., Neubig, G., and Lipton, Z. C · 2020
Later among the works it cites.
Explain your move: Understanding agent actions using specific and relevant feature attribution
Puri, N., Verma, S., Gupta, P., Kayastha, D., Deshmukh, S., Krishnamurthy, B., and Singh, S · 2020
Later among the works it cites.
Model agnostic multilevel explanations
Ramamurthy, K. N., Vinzamuri, B., Zhang, Y., and Dhurandhar, A · 2020
Later among the works it cites.
Search result explanations improve efficiency and trust
Ramos, J., and Eickhoff, C · 2020
Later among the works it cites.
Beyond individualized recourse: Interpretable and interactive summaries of actionable recourses
Rawal, K., and Lakkaraju, H · 2020
Later among the works it cites.
Interpretations are useful: Penalizing explanations to align neural networks with prior knowledge
Rieger, L., Singh, C., Murdoch, W. J., and Yu, B · 2020
Later among the works it cites.
Explainable inference on sequential data via memory-tracking
Rosa, B. L., Capobianco, R., and Nardi, D · 2020
Later among the works it cites.
Explainable machine learning for scientific insights and discoveries
Roscher, R., Bohn, B., Duarte, M. F., and Garcke, J · 2020
Later among the works it cites.
Evaluating attribution for graph neural networks
Sanchez-Lengeling, B., Wei, J., Lee, B., Reif, E., Wang, P., Qian, W., McCloskey, K., Colwell, L., and Wiltschko, A · 2020
Later among the works it cites.
Relation extraction with explanation
Shahbazi, H., Fern, X. Z., Ghaeini, R., and Tadepalli, P · 2020
Later among the works it cites.
Interpreting the latent space of gans for semantic face editing
Shen, Y., Gu, J., Tang, X., and Zhou, B · 2020
Later among the works it cites.
Dispersed exponential family mixture vaes for interpretable text generation
Shi, W., Zhou, H., Miao, N., and Li, L · 2020
Later among the works it cites.
Explanation by progressive exaggeration
Singla, S., Pollack, B., Chen, J., and Batmanghelich, K · 2020
Later among the works it cites.
When explanations lie: Why many modified BP attributions fail
Sixt, L., Granz, M., and Landgraf, T · 2020
Later among the works it cites.
Explainability fact sheets: a framework for systematic assessment of explainable approaches
Sokol, K., and Flach, P · 2020
Later among the works it cites.
Explanation perspectives from the cognitive sciences—a survey
Srinivasan, R., and Chander, A · 2020
Later among the works it cites.
Obtaining faithful interpretations from compositional neural networks
Subramanian, S., Bogin, B., Gupta, N., Wolfson, T., Singh, S., Berant, J., and Gardner, M · 2020
Later among the works it cites.
Dual learning for explainable recommendation: Towards unifying user preference prediction and review generation
Sun, P., Wu, L., Zhang, K., Fu, Y., Hong, R., and Wang, M · 2020
Later among the works it cites.
The many shapley values for model explanation
Sundararajan, M., and Najmi, A · 2020
Later among the works it cites.
Feature interaction interpretability: A case for explaining ad-recommendation systems via neural interaction detection
Tsang, M., Cheng, D., Liu, H., Feng, X., Zhou, E., and Liu, Y · 2020
Later among the works it cites.
How does this interaction affect me? interpretable attribution for feature interactions
Tsang, M., Rambhatla, S., and Liu, Y · 2020
Later among the works it cites.
Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics
Tseng, A., Shrikumar, A., and Kundaje, A · 2020
Later among the works it cites.
Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents
Tu, M., Huang, K., Wang, G., Huang, J., He, X., and Zhou, B · 2020
Later among the works it cites.
Unsupervised discovery of interpretable directions in the GAN latent space
Voynov, A., and Babenko, A · 2020
Later among the works it cites.
Pgm-explainer: Probabilistic graphical model explanations for graph neural networks
Vu, M. N., and Thai, M. T · 2020
Later among the works it cites.
SCOUT: self-aware discriminant counterfactual explanations
Wang, P., and Vasconcelos, N · 2020
Later among the works it cites.
Using small business banking data for explainable credit risk scoring
Wang, W., Lesner, C., Ran, A., Rukonic, M., Xue, J., and Shiu, E · 2020
Later among the works it cites.
Regional tree regularization for interpretability in deep neural networks
Wu, M., Parbhoo, S., Hughes, M. C., Kindle, R., Celi, L. A., Zazzi, M., Roth, V., and Doshi-Velez, F · 2020
Later among the works it cites.
Towards global explanations of convolutional neural networks with concept attribution
Wu, W., Su, Y., Chen, X., Zhao, S., King, I., Lyu, M. R., and Tai, Y · 2020
Later among the works it cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Wu, Z., Chen, Y., Kao, B., and Liu, Q · 2020
Later among the works it cites.
Explainable deep learning: A field guide for the uninitiated
Xie, N., Ras, G., van Gerven, M., and Doran, D · 2020
Later among the works it cites.
Explainable object-induced action decision for autonomous vehicles
Xu, Y., Yang, X., Gong, L., Lin, H., Wu, T., Li, Y., and Vasconcelos, N · 2020
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
Yeh, C., Kim, B., Arik, S. Ö., Li, C., Pfister, T., and Ravikumar, P · 2020
Later among the works it cites.
XGNN: towards model-level explanations of graph neural networks
Yuan, H., Tang, J., Hu, X., and Ji, S · 2020
Later among the works it cites.
Towards interpretable natural language understanding with explanations as latent variables
Zhou, W., Hu, J., Zhang, H., Liang, X., Sun, M., Xiong, C., and Tang, J · 2020
Later among the works it cites.
Towards a terminology for a fully contextualized xai
Bellucci, M., Delestre, N., Malandain, N., and Zanni-Merk, C · 2021
Later among the works it cites.
A survey on the explainability of supervised machine learning
Burkart, N., and Huber, M. F · 2021
Later among the works it cites.
The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment
Chicco, D., Warrens, M. J., and Jurman, G · 2021
Later among the works it cites.
Operationalizing human-centered perspectives in explainable ai
Ehsan, U., Wintersberger, P., Liao, Q. V., Mara, M., Streit, M., Wachter, S., Riener, A., and Riedl, M. O · 2021
Later among the works it cites.
SummEval: Re-evaluating Summarization Evaluation
Fabbri, A. R., KryÅ›ciÅ„ski, W., McCann, B., Xiong, C., Socher, R., and Radev, D · 2021
Later among the works it cites.
Evaluating local explanation methods on ground truth
Guidotti, R · 2021
Later among the works it cites.
The out-of-distribution problem in explainability and search methods for feature importance explanations
Hase, P., Xie, H., and Bansal, M · 2021
Later among the works it cites.
A review on explainability in multimodal deep neural nets
Joshi, G., Walambe, R., and Kotecha, K · 2021
Later among the works it cites.
Synthetic benchmarks for scientific research in explainable machine learning
Liu, Y., Khandagale, S., White, C., and Neiswanger, W · 2021
Later among the works it cites.
The role of explainability in creating trustworthy artificial intelligence for health care: A comprehensive survey of the terminology, design choices, and evaluation strategies
Markus, A. F., Kors, J. A., and Rijnbeek, P. R · 2021
Later among the works it cites.
A multidisciplinary survey and framework for design and evaluation of explainable ai systems
Mohseni, S., Zarei, N., and Ragan, E. D · 2021
Later among the works it cites.
Neural prototype trees for interpretable fine-grained image recognition
Nauta, M., van Bree, R., and Seifert, C · 2021
Later among the works it cites.
Embedding deep networks into visual explanations
Qi, Z., Khorram, S., and Fuxin, L · 2021
Later among the works it cites.
A survey of contrastive and counterfactual explanation generation methods for explainable artificial intelligence
Stepin, I., Alonso, J. M., Catala, A., and Pereira-Fariña, M · 2021
Later among the works it cites.
Notions of explainability and evaluation approaches for explainable artificial intelligence
Vilone, G., and Longo, L · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Evaluating the quality of machine learning explanations: A survey on methods and metrics
Zhou, J., Gandomi, A. H., Chen, F., and Holzinger, A · 2021
Later among the works it cites.
Interpretable machine learning: Moving from mythos to diagnostics
Chen, V., Li, J., Kim, J. S., Plumb, G., and Talwalkar, A · 2022
Closest in time.
A consistent and efficient evaluation strategy for attribution methods
Rong, Y., Leemann, T., Borisov, V., Kasneci, G., and Kasneci, E · 2022
Closest in time.