Fetching the paper…
Reading the bibliography…
Interpretability is the study of explaining models in understandable terms to humans.
On the (In)fidelity and Sensitivity of Explanations
Yeh, C.-K., Hsieh, C.-Y., Suggala, A., Inouye, D. I., Ravikumar, P. K., Sai Suggala, A., Inouye, D. I., and Ravikumar, P. K · 1901
Earlier work this paper cites.
Normlime: A new feature importance metric for explaining deep neural networks
Ahern, I., Noack, A., Guzmán-Nateras, L., Dou, D., Li, B., and Huan, J · 1909
Earlier work this paper cites.
One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques
Arya, V., Bellamy, R. K. E., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S. C., Houde, S., Liao, Q. V., Luss, R., Mojsilović, A., Mourad, S., Pedemonte, P., Raghavendra, R., Richards, J., Sattigeri, P., Shanmugam, K., Singh, M., Varshney, K. R., Wei, D., and Zhang, Y · 1909
Earlier work this paper cites.
Attention Interpretability Across NLP Tasks
Vashishth, S., Upadhyay, S., Tomar, G. S., and Faruqui, M · 1909
Earlier work this paper cites.
Did poincaré say “set theory is a disease”?
Gray, J · 1912
Earlier work this paper cites.
On the Legal Compatibility of Fairness Definitions
Xiang, A. and Raji, I. D · 1912
Earlier work this paper cites.
On Formally Undecidable Propositions of Principia Mathematica and Related Systems I
Gödel, K · 1931
Earlier work this paper cites.
Quantum Mechanics
Messiah, A · 1966
Earlier work this paper cites.
Georg Cantor and Pope Leo XIII: Mathematics, Theology, and the Infinite
Dauben, J. W · 1977
Earlier work this paper cites.
The Structure of Scientific Revolutions
Kuhn, T. S · 1996
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Classification by Set Cover: The Prototype Vector Machine
Bien, J. and Tibshirani, R · 2009
Earlier work this paper cites.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and Müller, K. R · 2010
Earlier work this paper cites.
The Bayesian case model: A generative approach for case-based reasoning and prototype classification
Kim, B., Rudin, C., and Shah, J · 2014
Earlier work this paper cites.
David Hilbert’s Radio Address
Smith, J · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K. H., and Bengio, Y · 2015
Earlier work this paper cites.
Visualizing and Understanding Recurrent Networks
Karpathy, A., Johnson, J., and Fei-Fei, L · 2015
Earlier work this paper cites.
Understanding Neural Networks Through Deep Visualization
Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H · 2015
Earlier work this paper cites.
Neural Module Networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
Layer-Wise Relevance Propagation for Neural Networks with Local Renormalization Layers
Binder, A., Montavon, G., Lapuschkin, S., Müller, K.-R., and Samek, W · 2016
Earlier work this paper cites.
Explainable Artificial Intelligence (XAI) DARPA-BAA-16-53
DARPA · 2016
Earlier work this paper cites.
Understanding Neural Networks through Representation Erasure
Li, J., Monroe, W., and Jurafsky, D · 2016
Earlier work this paper cites.
Nguyen, A., Yosinski, J., and Clune, J · 2016
Earlier work this paper cites.
”Why should i trust you?” Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Evaluating the Visualization of What a Deep Neural Network Has Learned
Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Muller, K.-R · 2016
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
European union regulations on algorithmic decision making and a ”right to explanation”
Goodman, B. and Flaxman, S · 2017
Earlier work this paper cites.
AI in the UK: Ready, Willing and Able?, 2017
House of Lords, U. G · 2017
Earlier work this paper cites.
Towards Faithful Model Explanation in NLP: A Survey
Lyu, Q., Apidianaki, M., and Callison-Burch, C · 2017
Earlier work this paper cites.
Feature Visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation
Chen, J., Song, L., Wainwright, M. J., and Jordan, M. I · 2018
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M · 2018
Earlier work this paper cites.
The mythos of model interpretability
Lipton, Z. C · 2018
Earlier work this paper cites.
Perturbation-Based Explanations of Prediction Models
Robnik-Šikonja, M. and Bohanec, M · 2018
Cited alongside, same era.
Noise-adding Methods of Saliency Map as Series of Higher Order Partial Derivative
Seo, J., Choe, J., Koo, J., Jeon, S., Kim, B., and Jeon, T · 2018
Cited alongside, same era.
Language Modeling Teaches You More than Translation Does: Lessons Learned Through Auxiliary Syntactic Task Analysis
Zhang, K. and Bowman, S · 2018
Cited alongside, same era.
A review of modularization techniques in artificial neural networks
Amer, M. and Maul, T · 2019
Cited alongside, same era.
Fairness and Machine Learning: Limitations and Opportunities
Barocas, S., Hardt, M., and Narayanan, A · 2019
Cited alongside, same era.
Analysis Methods in Neural Language Processing: A Survey
What we can’t measure, We can’t understand: Challenges to demographic data procurement in the pursuit of fairness
Andrus, M., Spitzer, E., Brown, J., and Xiang, A · 2021
Later among the works it cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Belinkov, Y · 2021
Later among the works it cites.
Exemplary Natural Images Explain Cnn Activations Better Than State-of-the-Art Feature Visualization
Borowski, J., Zimmermann, R. S., Schepers, J., Geirhos, R., Wallis, T. S., Bethge, M., and Brendel, W · 2021
Later among the works it cites.
A Mathematical Framework for Transformer Circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Later among the works it cites.
Companies Need More Workers. Why Do They Reject Millions of Résumés?
Fuller, J · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Belinkov, Y. and Glass, J · 2019
Cited alongside, same era.
Explainable Machine Learning in Deployment
Bhatt, U., Xiang, A., Sharma, S., Weller, A., Taly, A., Jia, Y., Ghosh, J., Puri, R., Moura, J. M. F., and Eckersley, P · 2019
Cited alongside, same era.
This looks like that: Deep learning for interpretable image recognition
Chen, C., Li, O., Tao, C., Barnett, A. J., Su, J., and Rudin, C · 2019
Cited alongside, same era.
What Does BERT Look at? An Analysis of BERT’s Attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Cited alongside, same era.
Visualizing and Measuring the Geometry of BERT
Coenen, A., Reif, E., Yuan, A., Kim, B., Pearce, A., Viégas, F., Wattenberg, M., Kim, A. Y. B., Pearce, A., Viégas, F., Wattenberg, M., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B · 2019
Cited alongside, same era.
ILIME: Local and Global Interpretable Model-Agnostic Explainer of Black-Box Decision
ElShawi, R., Sherif, Y., Al-Mallah, M., and Sakr, S · 2019
Cited alongside, same era.
Designing and Interpreting Probes with Control Tasks
Hewitt, J. and Liang, P · 2019
Cited alongside, same era.
Later among the works it cites.
Hidden Workers: Untapped Talent
Fuller, J. B., Ramen, M., Sage-gavin, E., and Hines, K · 2021
Later among the works it cites.
The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations
Hase, P., Xie, H., and Bansal, M · 2021
Later among the works it cites.
Have We Learned to Explain?: How Interpretability Methods Can Learn to Encode Predictions in their Interpretations
Jethani, N., Sudarshan, M., Aphinyanaphongs, Y., and Ranganath, R · 2021
Later among the works it cites.
Is Sparse Attention more Interpretable?
Meister, C., Lazov, S., Augenstein, I., and Cotterell, R · 2021
Later among the works it cites.
Rationales for Sequential Predictions
Vafa, K., Deng, Y., Blei, D., and Rush, A · 2021
Later among the works it cites.
How Well do Feature Visualizations Support Causal Understanding of CNN Activations?
Zimmermann, R. S., Borowski, J., Geirhos, R., Bethge, M., Wallis, T. S., and Brendel, W · 2021
Later among the works it cites.
“Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification
Bastings, J., Ebert, S., Zablotskaia, P., Sandholm, A., and Filippova, K · 2022
Later among the works it cites.
Local Structure Matters Most: Perturbation Study in NLU
Clouatre, L., Parthasarathi, P., Zouaq, A., and Chandar, S · 2022
Later among the works it cites.
Learning to Estimate Shapley Values with Vision Transformers
Covert, I. C., Kim, C., and Lee, S.-I · 2022
Later among the works it cites.
Machine learning in drug discovery: a review
Dara, S., Dhamercherla, S., Jadav, S. S., Babu, C. H. M., and Ahsan, M. J · 2022
Later among the works it cites.
Flora, M., Potvin, C., McGovern, A., and Handler, S · 2022
Later among the works it cites.
Han, T., Srinivas, S., and Lakkaraju, H · 2022
Later among the works it cites.
Fastshap: Real-Time Shapley Value Estimation
Jethani, N., Sudarshan, M., Covert, I., Lee, S. I., and Ranganath, R · 2022
Later among the works it cites.
Beyond interpretability: developing a language to shape our relationships with AI
Kim, B · 2022
Later among the works it cites.
The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective
Krishna, S., Han, T., Gu, A., Pombra, J., Jabbari, S., Wu, S., and Lakkaraju, H · 2022
Later among the works it cites.
Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining
Madsen, A., Meade, N., Adlakha, V., and Reddy, S · 2022
Later among the works it cites.
Bhalla, U., Srinivas, S., and Lakkaraju, H · 2023
Later among the works it cites.
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Chen, Y., Zhong, R., Ri, N., Zhao, C., He, H., Steinhardt, J., Yu, Z., and McKeown, K · 2023
Later among the works it cites.
Neural module networks: A review
Fashandi, H · 2023
Later among the works it cites.
Don’t trust your eyes: on the (un)reliability of feature visualizations
Geirhos, R., Zimmermann, R. S., Bilodeau, B., Brendel, W., and Kim, B · 2023
Later among the works it cites.
Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations
Huang, S., Mamidanna, S., Jangam, S., Zhou, Y., and Gilpin, L. H · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Later among the works it cites.
Measuring Faithfulness in Chain-of-Thought Reasoning
Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., Lukošiūtė, K., Nguyen, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., Larson, R., McCandlish, S., Kundu, S., Kadavath, S., Yang, S., Henighan, T., Maxwell, T., Telleen-Lawton, T., Hume, T., Hatfield-Dodds, Z., Kaplan, J., Brauner, J., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Meta · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
On Measuring Faithfulness of Natural Language Explanations
Parcalabescu, L. and Frank, A · 2023
Later among the works it cites.
Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
Schut, L., Tomasev, N., McGrath, T., Hassabis, D., Paquet, U., and Kim, B · 2023
Later among the works it cites.
Turpin, M., Michael, J., Perez, E., and Bowman, S. R · 2023
Later among the works it cites.
The Solvability of Interpretability Evaluation Metrics
Zhou, Y. and Shah, J · 2023
Later among the works it cites.
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
Agarwal, C., Tanneru, S. H., and Lakkaraju, H · 2024
Closest in time.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B · 2024
Closest in time.
Machine Learning Interpretability: A Survey on Methods and Metrics
Carvalho, D. V., Pereira, E. M., and Cardoso, J. S · 2079
Closest in time.