Fetching the paper…
Reading the bibliography…
Post-hoc explanation methods are gaining popularity for interpreting, understanding, and debugging neural networks.
Automating interpretability: Discovering and testing visual concepts learned by neural networks
Ghorbani, A.; Wexler, J.; and Kim, B. 2019 · 1902
Earlier work this paper cites.
Monte Carlo sampling methods using Markov chains and their applications
Hastings, W. K. 1970 · 1970
Earlier work this paper cites.
The mindlessness of ostensibly thoughtful action: The role of ”placebic” information in interpersonal interaction
Langer, E. J.; Blank, A.; and Chanowitz, B. 1978 · 1978
Earlier work this paper cites.
Trust in automation: Designing for appropriate reliance
Lee, J. D.; and See, K. A. 2004 · 2004
Earlier work this paper cites.
Getting a clue: A method for explaining uncertainty estimates
Antorán, J.; Bhatt, U.; Adel, T.; Weller, A.; and Hernández-Lobato, J. M. 2020 · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
Maaten, L. v. d.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D.; Bengio, Y.; Courville, A.; and Vincent, P. 2009 · 2009
Earlier work this paper cites.
On generating plausible counterfactual and semi-factual explanations for deep learning
Kenny, E. M.; and Keane, M. T. 2020 · 2009
Earlier work this paper cites.
MNIST handwritten digit database URL http://yann.lecun.com/exdb/mnist/
LeCun, Y.; and Cortes, C. 2010 · 2010
Earlier work this paper cites.
Handbook of markov chain monte carlo
Brooks, S.; Gelman, A.; Jones, G.; and Meng, X.-L. 2011 · 2011
Earlier work this paper cites.
MCMC using Hamiltonian dynamics
Neal, R. M.; et al. 2011 · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y.; Wang, T.; Coates, A.; Bissacco, A.; Wu, B.; and Ng, A. Y. 2011 · 2011
Earlier work this paper cites.
Evaluating Explanations: How much do explanations from the teacher aid students?
Pruthi, D.; Dhingra, B.; Soares, L. B.; Collins, M.; Lipton, Z. C.; Neubig, G.; and Cohen, W. W. 2020 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Simonyan, K.; Vedaldi, A.; and Zisserman, A. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; and Fergus, R. 2013 · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I.; Shlens, J.; and Szegedy, C. 2014 · 2014
Earlier work this paper cites.
The No-U-Turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
Hoffman, M. D.; and Gelman, A. 2014 · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
Zeiler, M. D.; and Fergus, R. 2014 · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, C.; Cornebise, J.; Kavukcuoglu, K.; and Wierstra, D. 2015 · 2015
Cited alongside, same era.
Measuring sample quality with Stein’s method
Gorham, J.; and Mackey, L. 2015 · 2015
Cited alongside, same era.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
Nguyen, A.; Yosinski, J.; and Clune, J. 2015 · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y.; and Ghahramani, Z. 2016 · 2016
Smoothgrad: removing noise by adding noise
Smilkov, D.; Thorat, N.; Kim, B.; Viégas, F.; and Wattenberg, M. 2017 · 2017
Later among the works it cites.
Adversarial discriminative domain adaptation
Tzeng, E.; Hoffman, J.; Saenko, K.; and Darrell, T. 2017 · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H.; Rasul, K.; and Vollgraf, R. 2017 · 2017
Later among the works it cites.
Sanity checks for saliency maps
Adebayo, J.; Gilmer, J.; Muelly, M.; Goodfellow, I.; Hardt, M.; and Kim, B. 2018 · 2018
Later among the works it cites.
Pyro: Deep Universal Probabilistic Programming
Bingham, E.; Chen, J. P.; Jankowiak, M.; Obermeyer, F.; Pradhan, N.; Karaletsos, T.; Singh, R.; Szerlip, P.; Horsfall, P.; and Goodman, N. D. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
”Why should I trust you?” Explaining the predictions of any classifier
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2016 · 2016
Cited alongside, same era.
Do GANs actually learn the distribution? An empirical study
Arora, S.; and Zhang, Y. 2017 · 2017
Cited alongside, same era.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D.; Zhou, B.; Khosla, A.; Oliva, A.; and Torralba, A. 2017 · 2017
Cited alongside, same era.
AIDE: An algorithm for measuring the accuracy of probabilistic inference algorithms
Cusumano-Towner, M.; and Mansinghka, V. K. 2017 · 2017
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F.; and Kim, B. 2017 · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Cited alongside, same era.
Not-So-CLEVR: learning same–different relations strains feedforward neural networks
Kim, J.; Ricci, M.; and Serre, T. 2018 · 2018
Later among the works it cites.
Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions
Li, O.; Liu, H.; Chen, C.; and Rudin, C. 2018 · 2018
Later among the works it cites.
The mythos of model interpretability
Lipton, Z. C. 2018 · 2018
Later among the works it cites.
Between-class learning for image classification
Tokozume, Y.; Ushiku, Y.; and Harada, T. 2018 · 2018
Later among the works it cites.
Generating Natural Adversarial Examples
Zhao, Z.; Dua, D.; and Singh, S. 2018 · 2018
Later among the works it cites.
This looks like that: deep learning for interpretable image recognition
Chen, C.; Li, O.; Tao, D.; Barnett, A.; Rudin, C.; and Su, J. K. 2019 · 2019
Later among the works it cites.
Scenic: a language for scenario specification and scene generation
Fremont, D. J.; Dreossi, T.; Ghosh, S.; Yue, X.; Sangiovanni-Vincentelli, A. L.; and Seshia, S. A. 2019 · 2019
Later among the works it cites.
The (un) reliability of saliency methods
Kindermans, P.-J.; Hooker, S.; Adebayo, J.; Alber, M.; Schütt, K. T.; Dähne, S.; Erhan, D.; and Kim, B. 2019 · 2019
Later among the works it cites.
Understanding neural networks via feature visualization: A survey
Nguyen, A.; Yosinski, J.; and Clune, J. 2019 · 2019
Later among the works it cites.
TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing
Odena, A.; Olsson, C.; Andersen, D.; and Goodfellow, I. 2019 · 2019
Later among the works it cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C. 2019 · 2019
Later among the works it cites.
On mixup training: Improved calibration and predictive uncertainty for deep neural networks
Thulasidasan, S.; Chennupati, G.; Bilmes, J. A.; Bhattacharya, T.; and Michalak, S. 2019 · 2019
Later among the works it cites.
How can we fool LIME and SHAP? Adversarial Attacks on Post hoc Explanation Methods
Slack, D.; Hilgard, S.; Jia, E.; Singh, S.; and Lakkaraju, H. 2020 · 2020
Closest in time.