Fetching the paper…
Reading the bibliography…
As black box explanations are increasingly being employed to establish model credibility in high-stakes settings, it is important to ensure that these explanations are accurate and reliable.
Locally weighted bayesian regression, January 1995
Andrew Moore · 1995
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Christopher M. Bishop · 2006
Earlier work this paper cites.
Regression
Ludwig Fahrmeir, Thomas Kneib, and Stefan Lang · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Active learning literature survey
Burr Settles · 2010
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Accurate intelligible models with pairwise interactions
Yin Lou, Rich Caruana, Johannes Gehrke, and Giles Hooker · 2013
Earlier work this paper cites.
Supersparse linear integer models for interpretable classification
Berk Ustun, Stefano Traca, and Cynthia Rudin · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L. Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
The bayesian case model: A generative approach for case-based reasoning and prototype classification
Been Kim, Cynthia Rudin, and Julie Shah · 2015
Earlier work this paper cites.
Mapping chemical performance on molecular structures using locally interpretable explanations
Leanne S Whitmore, Anthe George, and Corey M Hudson · 2016
Earlier work this paper cites.
Machine bias
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner · 2016
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
Himabindu Lakkaraju, Stephen H Bach, and Jure Leskovec · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Interpretability via model extraction
Osbert Bastani, Carolyn Kim, and Hamsa Bastani · 2017
Cited alongside, same era.
Uci machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Learning certifiably optimal rule lists for categorical data
Elaine Angelino, Nicholas Larus-Stone, Daniel Alabi, Margo Seltzer, and Cynthia Rudin · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Sandra Wachter, Brent Mittelstadt, and Chris Russell · 2017
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon · 2019
Later among the works it cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Later among the works it cites.
Muhammad Rehman Zafar and Naimul Mefraz Khan · 2019
Later among the works it cites.
On the (in) fidelity and sensitivity of explanations
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar · 2019
Later among the works it cites.
Can i trust the explainer? verifying post-hoc explanatory methods
Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model agnostic supervised local explanations
Gregory Plumb, Denali Molitor, and Ameet S Talwalkar · 2018
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Cited alongside, same era.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S. Jaakkola · 2018
Cited alongside, same era.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach · 2018
Cited alongside, same era.
A symbolic approach to explaining bayesian network classifiers
Andy Shih, Arthur Choi, and Adnan Darwiche · 2018
Cited alongside, same era.
Explaining deep learning models – a bayesian non-parametric approach
Wenbo Guo, Sui Huang, Yunzhe Tao, Xinyu Xing, and Lin Lin · 2018
Cited alongside, same era.
Later among the works it cites.
Bim: Towards quantitative evaluation of interpretability methods with ground truth
Mengjiao Yang and Been Kim · 2019
Later among the works it cites.
Gamut: A design probe to understand how data scientists understand machine learning models
Fred Hohman, Andrew Head, Rich Caruana, Robert DeLine, and Steven M Drucker · 2019
Later among the works it cites.
Certifiably robust interpretation in deep learning
Alexander Levine, Sahil Singla, and Soheil Feizi · 2019
Later among the works it cites.
Abduction-based explanations for machine learning models
Alexey Ignatiev, Nina Narodytska, and Joao Marques-Silva · 2019
Later among the works it cites.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Closest in time.
Gradient-based Analysis of NLP Models is Manipulable
Junlin Wang, Jens Tuyls, Eric Wallace, and Sameer Singh · 2020
Closest in time.
Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan · 2020
Closest in time.
" how do i fool you?" manipulating user trust via misleading black box explanations
Himabindu Lakkaraju and Osbert Bastani · 2020
Closest in time.
Damien Garreau and Ulrike von Luxburg · 2020
Closest in time.
Concise explanations of neural networks using adversarial training
Prasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu, and Somesh Jha · 2020
Closest in time.
On tractable representations of binary neural networks
Weijia Shi, Andy Shih, Adnan Darwiche, and Arthur Choi · 2020
Closest in time.
Baylime: Bayesian local interpretable model-agnostic explanations
Xingyu Zhao, Xiaowei Huang, Valentin Robu, and David Flynn · 2020
Closest in time.
How much can i trust you? – quantifying uncertainties in explaining neural networks
Kirill Bykov, Marina Höhne, Klaus-Robert Müller, Shinichi Nakajima, and Marius Kloft · 2020
Closest in time.
Improving kernelshap: Practical shapley value estimation using linear regression
Ian Covert and Su-In Lee · 2021
Closest in time.
Probabilistic sufficient explanations
Eric Wang, Pasha Khosravi, and Guy Van den Broeck · 2021
Closest in time.
Towards the unification and robustness of perturbation and gradient based explanations
Sushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay, Steven Wu, and Himabindu Lakkaraju · 2021
Closest in time.