Fetching the paper…
Reading the bibliography…
Methods for interpreting machine learning black-box models increase the outcomes' transparency and in turn generates insight into the reliability and fairness of the algorithms.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and Muller, K.-R · 2010
Earlier work this paper cites.
Interpretable classifiers using rules and bayesian analysis: Building a better stroke prediction model
Letham, B., Rudin, C., McCormick, T. H., and Madigan, D · 2015
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
Lakkaraju, H., Bach, S. H., and Leskovec, J · 2016
Earlier work this paper cites.
How We Analyzed the COMPAS Recidivism Algorithm, 2016
Larson, J. L., Mattu, S., Kirchner, L., and Angwin, J · 2016
Earlier work this paper cites.
The mythos of model interpretability
Lipton, Z. C · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Interpretable classification models for recidivism prediction
Zeng, J., Ustun, B., and Rudin, C · 2017
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Ribeiro, M. T., Singh, S., and Guestrin, C
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Later among the works it cites.
On the robustness of interpretability methods
Alvarez-Melis, D. and Jaakkola, T. S · 2018
Later among the works it cites.
Scalable and accurate deep learning with electronic health records
Rajkomar, A., Oren, E., Chen, K., Dai, A. M., Hajaj, N., Hardt, M., Liu, P. J., Liu, X., Marcus, J., Sun, M., et al · 2018
Later among the works it cites.
Interpretation of neural networks is fragile
Ghorbani, A., Abid, A., and Zou, J · 2019
Closest in time.
Developing the sensitivity of lime for better machine learning explanation
Lee, E., Braines, D., Stiffler, M., Hudler, A., and Harborne, D · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited in the paper.
Semantically equivalent adversarial rules for debugging nlp models
Ribeiro, M. T., Singh, S., and Guestrin, C
Cited in the paper.
Learning global additive explanations for neural nets using model distillation
Tan, S., Caruana, R., Hooker, G., Koch, P., and Gordo, A
Cited in the paper.
Distill-and-compare: Auditing black-box models using transparent model distillation
Tan, S., Caruana, R., Hooker, G., and Lou, Y
Cited in the paper.