Fetching the paper…
Reading the bibliography…
A feature-based model explanation denotes how much each input feature contributes to a model's output for a given data point.
A value for n-person games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The doctor–patient relationship: challenges, opportunities, and strategies
Susan Dorr Goold and Mack Lipkin Jr · 1999
Earlier work this paper cites.
Inference to the best explanation
Peter Lipton · 2003
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Muller · 2010
Earlier work this paper cites.
Statistical decision theory and Bayesian analysis
James O Berger · 2013
Earlier work this paper cites.
Minimal model explanations
Robert W Batterman and Collin C Rice · 2014
Earlier work this paper cites.
Explaining prediction models and individual predictions with feature contributions
Erik Štrumbelj and Igor Kononenko · 2014
Earlier work this paper cites.
Explaining explanation
David-Hillel Ruben · 2015
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller · 2016
Earlier work this paper cites.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Oztireli, and Markus Gross · 2018
Cited alongside, same era.
What do different evaluation metrics tell us about saliency models?
Zoya Bylinskii, Tilke Judd, Aude Oliva, Antonio Torralba, and Frédo Durand · 2018
Cited alongside, same era.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin Wainwright, and Michael Jordan · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Cited alongside, same era.
Shedding Light on Black Box Algorithms
Milo Honegger · 2018
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks
Christopher J Hazard, Christopher Fusting, Michael Resnick, Michael Auerbach, Michael Meehan, and Valeri Korobov · 2019
Later among the works it cites.
Ted: Teaching ai to explain its decisions
Michael Hind, Dennis Wei, Murray Campbell, Noel CF Codella, Amit Dhurandhar, Aleksandra Mojsilović, Karthikeyan Natesan Ramamurthy, and Kush R Varshney · 2019
Later among the works it cites.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Later among the works it cites.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Later among the works it cites.
Human evaluation of models built for interpretability
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J Gershman, and Finale Doshi-Velez · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Alvarez Melis and Tommi Jaakkola · 2018
Cited alongside, same era.
Methods for interpreting and understanding deep neural networks
Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller · 2018
Cited alongside, same era.
Model agnostic supervised local explanations
Gregory Plumb, Denali Molitor, and Ameet S Talwalkar · 2018
Cited alongside, same era.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach · 2018
Cited alongside, same era.
Benchmarking deep learning models on large healthcare datasets
Sanjay Purushotham, Chuizheng Meng, Zhengping Che, and Yan Liu · 2018
Cited alongside, same era.
Building human-machine trust via interpretability
Umang Bhatt, Pradeep Ravikumar, et al · 2019
Cited alongside, same era.
Towards aggregating weighted feature attributions
Umang Bhatt, Pradeep Ravikumar, and José M. F. Moura · 2019
Cited alongside, same era.
Evaluating explanation methods for deep learning in security
Alexander Warnecke, Daniel Arp, Christian Wressnegger, and Konrad Rieck · 2019
Later among the works it cites.
BIM: Towards quantitative evaluation of interpretability methods with ground truth
Mengjiao Yang and Been Kim · 2019
Later among the works it cites.
Evaluating explanation without ground truth in interpretable machine learning
Fan Yang, Mengnan Du, and Xia Hu · 2019
Later among the works it cites.
On the (in) fidelity and sensitivity of explanations
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar · 2019
Later among the works it cites.
Towards a unified evaluation of explanation methods without ground truth
Hao Zhang, Jiayi Chen, Haotian Xue, and Quanshi Zhang · 2019
Later among the works it cites.
Explainable machine learning in deployment
Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José M. F. Moura, and Peter Eckersley · 2020
Closest in time.
On network science and mutual information for explaining deep neural networks
B. Davis, U. Bhatt, K. Bhardwaj, R. Marculescu, and J. M. F. Moura · 2020
Closest in time.
Measuring and improving the quality of visual explanations
Agnieszka Grabska-Barwińska · 2020
Closest in time.
Towards ground truth evaluation of visual explanations
Ahmed Osman, Leila Arras, and Wojciech Samek · 2020
Closest in time.
Irof: a low resource evaluation metric for explanation methods
Laura Rieger and Lars Kai Hansen · 2020
Closest in time.
Interpreting interpretations: Organizing attribution methods by criteria
Zifan Wang, Piotr Mardziel, Anupam Datta, and Matt Fredrikson · 2020
Closest in time.