Fetching the paper…
Reading the bibliography…
How can we understand classification decisions made by deep neural networks? Many existing explainability methods rely solely on correlations and fail to account for confounding, which may result in potentially misleading explanations.
Causes and explanations: A structural-model approach. part ii: Explanations
Halpern, J. Y. and Pearl, J · 2005
Earlier work this paper cites.
Making things happen: A theory of causal explanation
Woodward, J · 2005
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Deng, L · 2012
Earlier work this paper cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J · 2012
Earlier work this paper cites.
The bayesian case model: A generative approach for case-based reasoning and prototype classification
Kim, B., Rudin, C., and Shah, J. A · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Methods and models for interpretable linear classification
Ustun, B. and Rudin, C · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Sohn, K., Lee, H., and Yan, X · 2015
Cited alongside, same era.
Model-agnostic interpretability of machine learning
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Cited alongside, same era.
Real time image saliency for black box classifiers
Dabkowski, P. and Gal, Y · 2017
Cited alongside, same era.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. C. and Vedaldi, A · 2017
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Kim, B. and Doshi-Velez · 2017
Cited alongside, same era.
Causalgan: Learning causal implicit generative models with adversarial training
Explaining image classifiers by counterfactual generation
Chang, C.-H., Creager, E., Goldenberg, A., and Duvenaud, D · 2018
Later among the works it cites.
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation
Choi, Y., Choi, M., Kim, M., Ha, J.-W., Kim, S., and Choo, J · 2018
Later among the works it cites.
Generating counterfactual explanations with natural language
Hendricks, L. A., Hu, R., Darrell, T., and Akata, Z · 2018
Later among the works it cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., and Sayres, R · 2018
Later among the works it cites.
Counterfactual visual explanations
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., and Lee, S · 2019
Closest in time.
Learning not to learn: Training deep neural networks with biased data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kocaoglu, M., Snyder, C., Dimakis, A. G., and Vishwanath, S · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M · 2017
Cited alongside, same era.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A · 2017
Cited alongside, same era.
Kim, B., Kim, H., Kim, K., Kim, S., and Kim, J · 2019
Closest in time.
Direct optimization through arg \arg max \max for discrete variational auto-encoder
Lorberbom, G., Gane, A., Jaakkola, T., and Hazan, T · 2019
Closest in time.
Towards Quantitative Evaluation of Interpretability Methods with Ground Truth
Yang, M. and Kim, B · 2019
Closest in time.