Fetching the paper…
Reading the bibliography…
Interpretability techniques aim to provide the rationale behind a model's decision, typically by explaining either an individual prediction (local explanation, e.g.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Hui Zou and Trevor Hastie · 2005
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Frank-wolfe bayesian quadrature: Probabilistic integration with theoretical guarantees
François-Xavier Briol, Chris Oates, Mark Girolami, and Michael A Osborne · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate, 2016
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
"Why should i trust you?" Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje · 2016
Earlier work this paper cites.
Learning certifiably optimal rule lists for categorical data
Elaine Angelino, Nicholas Larus-Stone, Daniel Alabi, Margo Seltzer, and Cynthia Rudin · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Been Doshi-Velez, Finale; Kim · 2017
Earlier work this paper cites.
Interpretation of neural networks is fragile, 2017
Amirata Ghorbani, Abubakar Abid, and James Zou · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Fast estimation of tr(f(a)) via stochastic lanczos quadrature
Shashanka Ubaru, Jie Chen, and Yousef Saad · 2017
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for Deep Neural Networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Cited alongside, same era.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S. Jaakkola · 2018
Cited alongside, same era.
Considerations for Evaluation and Generalization in Interpretable Machine Learning
Finale Doshi-Velez and Been Kim · 2018
Cited alongside, same era.
Explainable and Interpretable Models in Computer Vision and Machine Learning
Hugo Jair Escalante, Sergio Escalera, Isabelle Guyon, Xavier Baró, Yağmur Güçlütürk, Umut Güçlü, and Marcel van Gerven · 2018
Cited alongside, same era.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Later among the works it cites.
When explanations lie: Why modified bp attribution fails
Leon Sixt, Maximilian Granz, and Tim Landgraf · 2019
Later among the works it cites.
Using a Deep Learning Algorithm and Integrated Gradients Explanation to Assist Grading for Diabetic Retinopathy
Rory Sayres, Ankur Taly, Ehsan Rahimy, Katy Blumer, David Coz, Naama Hammel, Jonathan Krause, Arunachalam Narayanaswamy, Zahra Rastegar, Derek Wu, Shawn Xu, Scott Barb, Anthony Joseph, Michael Shumski, Jesse Smith, Arjun B Sood, Greg S Corrado, Lily Peng, and Dale R Webster · 2019
Later among the works it cites.
What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use
Sana Tonekaboni, Shalmali Joshi, Melissa D McCradden, and Anna Goldenberg · 2019
Later among the works it cites.
Global aggregations of local explanations for black box models
Ilse van der Linden, Hinda Haned, and Evangelos Kanoulas · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regression concept vectors for bidirectional explanations in histopathology
Mara Graziani, Vincent Andrearczyk, and Henning Müller · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C. Lipton · 2018
Cited alongside, same era.
Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Places: A 10 million image database for scene recognition
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba · 2018
Cited alongside, same era.
Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Values Approximation
Marco Ancona, Cengiz Öztireli, and Markus Gross · 2019
Cited alongside, same era.
This looks like that: Deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Cited alongside, same era.
Later among the works it cites.
On the (in)fidelity and sensitivity of explanations
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar · 2019
Later among the works it cites.
BIM: towards quantitative evaluation of interpretability methods with ground truth
Mengjiao Yang and Been Kim · 2019
Later among the works it cites.
Debiasing concept bottleneck models with instrumental variables
Mohammad Taha Bahadori and David E. Heckerman · 2020
Later among the works it cites.
Explainable machine learning in deployment
Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José M F Moura, and Peter Eckersley · 2020
Later among the works it cites.
Tensorflow official model garden, 2020
Chen Chen, Xianzhi Du, Le Hou, Jaeyoun Kim, Jing Li, Yeqing Li, Abdullah Rashwan, Fan Yang, and Hongkun Yu · 2020
Later among the works it cites.
Improving performance of deep learning models with axiomatic attribution priors and expected gradients, 2020
Gabriel Erion, Joseph D. Janizek, Pascal Sturmfels, Scott Lundberg, and Su-In Lee · 2020
Later among the works it cites.
Understanding integrated gradients with smoothtaylor for deep neural network attribution
Gary S. W. Goh, Sebastian Lapuschkin, Leander Weber, Wojciech Samek, and Alexander Binder · 2020
Later among the works it cites.
Summit: Scaling deep learning interpretability by visualizing activation and attribution summarizations
Fred Hohman, Haekyu Park, Caleb Robinson, and Duen Horng Polo Chau · 2020
Later among the works it cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Later among the works it cites.
An ontology-based approach to black-box sequential data classification explanations
Cecilia Panigutti, Alan Perotti, Dino Pedreschi, and Dino 2020 Pedreschi · 2020
Later among the works it cites.
On the interpretability of artificial intelligence in radiology: Challenges and opportunities
Mauricio Reyes, Raphael Meier, Sérgio Pereira, Carlos A Silva, Fried-Michael Dahlweid, Hendrik von Tengg-Kobligk, Ronald M Summers, and Roland Wiest · 2020
Later among the works it cites.
Visualizing the Impact of Feature Attribution Baselines
Pascal Sturmfels, Scott Lundberg, and Su-In Lee · 2020
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
Mingxing Tan and Quoc V. Le · 2020
Later among the works it cites.
Towards global explanations of convolutional neural networks with concept attribution
Weibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao, Irwin King, Michael R. Lyu, and Yu-Wing Tai · 2020
Later among the works it cites.
On completeness-aware Concept-Based explanations in deep neural networks
Chih-Kuan Yeh, Been Kim, Sercan Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar · 2020
Later among the works it cites.
What I cannot predict, I do not understand: A Human-Centered evaluation framework for explainability methods
Thomas Fel, Julien Colin, Remi Cadene, and Thomas Serre · 2021
Closest in time.
The false hope of current approaches to explainable artificial intelligence in health care
Marzyeh Ghassemi, Luke Oakden-Rayner, and Andrew L Beam · 2021
Closest in time.
Guided integrated gradients: An adaptive path method for removing noise
Andrei Kapishnikov, Subhashini Venugopalan, Besim Avci, Ben Wedin, Michael Terry, and Tolga Bolukbasi · 2021
Closest in time.
Concept-Based Model Explanations for Electronic Health Records
Diana Mincu, Eric Loreaux, Shaobo Hou, Sebastien Baur, Ivan Protsyuk, Martin Seneviratne, Anne Mottram, Nenad Tomasev, Alan Karthikesalingam, and Jessica Schrouff · 2021
Closest in time.