Fetching the paper…
Reading the bibliography…
Being able to provide explanations for a model's decision has become a central requirement for the development, deployment, and adoption of machine learning models.
The Direction of Time
Reichenbach, H · 1956
Earlier work this paper cites.
Classifying the classifier: dissecting the weight space of neural networks
Eilertsen, G., Jönsson, D., Ropinski, T., Unger, J., and Ynnerman, A · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. J · 2002
Earlier work this paper cites.
Causal inference using potential outcomes: Design, modeling, decisions
Rubin, D. B · 2005
Earlier work this paper cites.
Rewriting a deep generative model
Bau, D., Liu, S., Wang, T., Zhu, J., and Torralba, A · 2007
Earlier work this paper cites.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and Müller, K.-R · 2009
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
General data protection regulation
Parliament and of the European Union, C · 2016
Earlier work this paper cites.
Grad-cam: Why did you say that?
Selvaraju, R. R., Das, A., Vedantam, R., Cogswell, M., Parikh, D., and Batra, D · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Wachter, S., Mittelstadt, B., and Russell, C · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Cited alongside, same era.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
Rieger, L., Singh, C., Murdoch, W. J., and Yu, B · 2020
Later among the works it cites.
Predicting neural network accuracy from weights, 2020
Unterthiner, T., Keysers, D., Gelly, S., Bousquet, O., and Tolstikhin, I · 2020
Later among the works it cites.
Attribution in scale and space
Xu, S., Venugopalan, S., and Sundararajan, M · 2020
Later among the works it cites.
Guided integrated gradients: An adaptive path method for removing noise
Kapishnikov, A., Venugopalan, S., Avci, B., Wedin, B., Terry, M., and Bolukbasi, T · 2021
Later among the works it cites.
Counterfactual mean embeddings
Muandet, K., Kanagawa, M., Saengkyongam, S., and Marukatat, S · 2021
Later among the works it cites.
Conditional distributional treatment effect with kernel conditional mean embeddings and U-statistic regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nie, W., Zhang, Y., and Patel, A · 2018
Cited alongside, same era.
Manipulating and measuring model interpretability
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Vaughan, J. W., and Wallach, H · 2018
Cited alongside, same era.
Predicting the generalization gap in deep networks with margin distributions
Jiang, Y., Krishnan, D., Mobahi, H., and Bengio, S · 2019
Cited alongside, same era.
The (un) reliability of saliency methods
Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., and Kim, B · 2019
Cited alongside, same era.
Evaluating saliency map explanations for convolutional neural networks: a user study
Alqaraawi, A., Schuessler, M., Weiß, P., Costanza, E., and Berthouze, N · 2020
Cited alongside, same era.
Are visual explanations useful? a case study in model-in-the-loop prediction
Chu, E., Roy, D., and Andreas, J · 2020
Cited alongside, same era.
Underspecification presents challenges for credibility in modern machine learning
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al · 2020
Cited alongside, same era.
Park, J., Shalit, U., Schölkopf, B., and Muandet, K · 2021
Later among the works it cites.
Rethinking the role of gradient-based attribution methods for model interpretability
Srinivas, S. and Fleuret, F · 2021
Later among the works it cites.
Robust models are more interpretable because attributions look normal
Wang, Z., Fredrikson, M., and Datta, A · 2021
Later among the works it cites.
Causal interpretations of black-box models
Zhao, Q. and Hastie, T · 2021
Later among the works it cites.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Adebayo, J., Muelly, M., Abelson, H., and Kim, B · 2022
Closest in time.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B · 2022
Closest in time.
Locating and editing factual knowledge in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Closest in time.
Direct and indirect effects
Pearl, J · 2022
Closest in time.