Fetching the paper…
Reading the bibliography…
The ability to explain decisions made by AI systems is highly sought after, especially in domains where human lives are at stake such as medicine or autonomous vehicles.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 1903
Earlier work this paper cites.
An introduction to cybernetics
Ashby, W. R · 1956
Earlier work this paper cites.
An informational measure of correlation
Linfoot, E · 1957
Earlier work this paper cites.
XPLAIN: a system for creating and explaining expert consulting programs
Swartout, W. R · 1983
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
Opper, M., Kinzel, W., Kleinz, J., and Nehl, R · 1990
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
Breiman, L · 2001
Earlier work this paper cites.
Feature extraction by non-parametric mutual information maximization
Torkkola, K · 2003
Earlier work this paper cites.
Current status of methods for defining the applicability domain of (quantitative) structure-activity relationships
Netzeva, T. I., Worth, A. P., Aldenberg, T., Benigni, R., Cronin, M. T., Gramatica, P., Jaworska, J. S., Kahn, S., Klopman, G., Marchant, C. A., Myatt, G., Nikolova-Jeliazkova, N., Patlewicz, G. Y., Perkins, R., Roberts, D. W., Schultz, T. W., Stanton, D. T., van de Sandt, J. J., Tong, W., Veith, G., and Yang, C · 2005
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Comparison of different approaches to define the applicability domain of QSAR models
Sahigara, F., Mansouri, K., Ballabio, D., Mauri, A., Consonni, V., and Todeschini, R · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Principles of explanatory debugging to personalize interactive machine learning
Kulesza, T., Burnett, M., Wong, W.-K., and Stumpf, S · 2015
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. A · 2015
Earlier work this paper cites.
Understanding neural networks through deep visualization
Yosinski, J., Clune, J., Nguyen, A. M., Fuchs, T. J., and Lipson, H · 2015
Earlier work this paper cites.
Inverting visual representations with convolutional networks
Dosovitskiy, A. and Brox, T · 2016
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
The mythos of model interpretability
Lipton, Z. C · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Nguyen, A. M., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J · 2016
Earlier work this paper cites.
Why should I trust you?
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. and Vedaldi, A · 2017
Earlier work this paper cites.
Distilling a neural network into a soft decision tree
Frosst, N. and Hinton, G · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Kendall, A. and Gal, Y · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S · 2017
Cited alongside, same era.
Explaining nonlinear classification decisions with deep taylor decomposition
Montavon, G., Lapuschkin, S., Binder, A., Samek, W., and Müller, K.-R · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Deep ReLU networks have surprisingly few activation patterns
Hanin, B. and Rolnick, D · 2019
Later among the works it cites.
A benchmark for interpretability methods in deep neural networks
Hooker, S., Erhan, D., Kindermans, P., and Kim, B · 2019
Later among the works it cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Later among the works it cites.
Interpreting black box predictions using fisher kernels
Khanna, R., Kim, B., Ghosh, J., and Koyejo, S · 2019
Later among the works it cites.
The (un)reliability of saliency methods
Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., and Kim, B · 2019
Later among the works it cites.
Encoding Visual Attributes in Capsules for Explainable Medical Diagnoses
LaLonde, R., Torigian, D., and Bagci, U · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Cited alongside, same era.
Interpretable machine learning in healthcare
Ahmad, M. A., Eckert, C., and Teredesai, A · 2018
Cited alongside, same era.
Hybrid strategies towards safe self-aware superintelligent systems
Aliman, N.-M. and Kester, L · 2018
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D. and Jaakkola, T. S · 2018
Cited alongside, same era.
Machine learning of energetic material properties
Barnes, B. C., Elton, D. C., Boukouvalas, Z., Taylor, D. E., Mattson, W. D., Fuge, M. D., and Chung, P. W · 2018
Cited alongside, same era.
Iteratively unveiling new regions of interest in deep learning models
Bordes, F., Berthier, T., Jorio, L. D., Vincent, P., and Bengio, Y · 2018
Cited alongside, same era.
Later among the works it cites.
Relevance in the eye of the beholder: Diagnosing classifications based on visualised layerwise relevance propagation
Lie, C · 2019
Later among the works it cites.
What does it mean to understand a neural network?
Lillicrap, T. P. and Kording, K. P · 2019
Later among the works it cites.
Knowing what you know in brain segmentation using Bayesian deep neural networks
McClure, P., Rho, N., Lee, J. A., Kaczmarzyk, J. R., Zheng, C. Y., Ghosh, S. S., Nielson, D. M., Thomas, A. G., Bandettini, P., and Pereira, F · 2019
Later among the works it cites.
Müller, H. and Holzinger, A · 2019
Later among the works it cites.
Definitions, methods, and applications in interpretable machine learning
Murdoch, W. J., Singh, C., Kumbier, K., Abbasi-Asl, R., and Yu, B · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
A deep learning framework for neuroscience
Richards, B. A., Lillicrap, T. P., Beaudoin, P., Bengio, Y., Bogacz, R., Christensen, A., Clopath, C., Costa, R. P., de Berker, A., Ganguli, S., Gillon, C. J., Hafner, D., Kepecs, A., Kriegeskorte, N., Latham, P., Lindsay, G. W., Miller, K. D., Naud, R., Pack, C. C., Poirazi, P., Roelfsema, P., Sacramento, J., Saxe, A., Scellier, B., Schapiro, A. C., Senn, W., Wayne, G., Yamins, D., Zenke, F., Zylberberg, J., Therien, D., and Kording, K. P · 2019
Later among the works it cites.
Identifying weights and architectures of unknown relu networks
Rolnick, D. and Kording, K. P · 2019
Later among the works it cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C · 2019
Later among the works it cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2019
Later among the works it cites.
An interpretable deep hierarchical semantic convolutional neural network for lung nodule malignancy classification
Shen, S., Han, S. X., Aberle, D. R., Bui, A. A., and Hsu, W · 2019
Later among the works it cites.
A jamming transition from under- to over-parametrization affects generalization in deep learning
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2019
Later among the works it cites.
On the (in)fidelity and sensitivity for explanations
Yeh, C.-K., Hsieh, C.-Y., Suggala, A. S., Inouye, D. I., and Ravikumar, P · 2019
Later among the works it cites.
Interpreting deep visual representations via network dissection
Zhou, B., Bau, D., Oliva, A., and Torralba, A · 2019
Later among the works it cites.
Accurately identifying vertebral levels in large datasets
Elton, D., Sandfort, V., Pickhardt, P. J., and Summers, R. M · 2020
Closest in time.
Neuron shapley: Discovering the responsible neurons
Ghorbani, A. and Zou, J · 2020
Closest in time.
Evaluating explainable ai: Which algorithmic explanations help users predict model behavior?
Hase, P. and Bansal, M · 2020
Closest in time.
Direct fit to nature: An evolutionary perspective on biological and artificial neural networks
Hasson, U., Nastase, S. A., and Goldstein, A · 2020
Closest in time.
Visualization approach to assess the robustness of neural networks for medical image classification
Sutre, E. T., Colliot, O., Dormont, D., and Burgos, N · 2020
Closest in time.