Fetching the paper…
Reading the bibliography…
Explainable machine learning offers the potential to provide stakeholders with insights into model behavior by using various methods such as feature importance scores, counterfactual explanations, or influential training data.
"Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
Yujia Zhang, Kuangyan Song, Yiming Sun, Sarah Tan, and Madeleine Udell. 2019 · 1904
Earlier work this paper cites.
Adversarial Examples Are Not Bugs, They Are Features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. 2019 · 1905
Earlier work this paper cites.
A Value for n-Person Games
Lloyd S Shapley. 1953 · 1953
Earlier work this paper cites.
Detection of influential observation in linear regression
R Dennis Cook. 1977 · 1977
Earlier work this paper cites.
Statistics and Causal Inference
Paul W. Holland. 1986 · 1986
Earlier work this paper cites.
Causality: models, reasoning and inference . Vol. 29
Judea Pearl. 2000 · 2000
Earlier work this paper cites.
Data squashing: constructing summary data sets
William DuMouchel. 2002 · 2002
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert MÞller. 2010 · 2010
Earlier work this paper cites.
Supervisory Guidance on Model Risk Management
Board of Governors of the Federal Reserve System. 2011 · 2011
Earlier work this paper cites.
Explaining prediction models and individual predictions with feature contributions
Erik Štrumbelj and Igor Kononenko. 2014 · 2014
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Examples are not Enough, Learn to Criticize! Criticism for Interpretability. In Advances in Neural Information Processing Systems
Rajiv Khanna Been Kim and Sanmi Koyejo. 2016 · 2016
Earlier work this paper cites.
JB Heaton, Nicholas G Polson, and Jan Hendrik Witte. 2016 · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . ACM, 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Stealing machine learning models via prediction apis. In 25th { \{ USENIX } \} Security Symposium ( { \{ USENIX } \} Security 16) . 601–618
Florian Tramèr, Fan Zhang, Ari Juels, Michael K Reiter, and Thomas Ristenpart. 2016 · 2016
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Interpretable Explanations of Black Boxes by Meaningful Perturbation
Ruth Fong and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 (ICML 2017) . Journal of Machine Learning Research, 1885–1894
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017 · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2017 · 2017
Earlier work this paper cites.
Right for the right reasons: training differentiable models by constraining their explanations. In Proceedings of the 26th International Joint Conference on Artificial Intelligence . AAAI Press, 2662–2670
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 (ICML 2017) . Journal of Machine Learning Research, 3145–3153
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 (ICML 2017) . Journal of Machine Learning Research, 3319–3328
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GPDR
Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2017 · 2017
Cited alongside, same era.
Discovering interpretable representations for both deep generative and discriminative models. In International Conference on Machine Learning . 50–59
Tameem Adel, Zoubin Ghahramani, and Adrian Weller. 2018 · 2018
Cited alongside, same era.
Excitation backprop for RNNs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1440–1449
Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Value Approximation. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.), Vol. 97. PMLR, Long Beach, California, USA, 272–281
Marco Ancona, Cengiz Oztireli, and Markus Gross. 2019 · 2019
Closest in time.
Towards aggregating weighted feature attributions
Umang Bhatt, Pradeep Ravikumar, and José MF Moura. 2019 · 2019
Closest in time.
Neural Network Attributions: A Causal Perspective. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.), Vol. 97. PMLR, Long Beach, California, USA, 981–990
Aditya Chattopadhyay, Piyushi Manupriya, Anirban Sarkar, and Vineeth N Balasubramanian. 2019 · 2019
Closest in time.
L-shapley and c-shapley: Efficient model interpretation for structured data
Jianbo Chen, Le Song, Martin J Wainwright, and Michael I Jordan. [n. d.] · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarah Adel Bargal, Andrea Zunino, Donghyun Kim, Jianming Zhang, Vittorio Murino, and Stan Sclaroff. 2018 · 2018
Cited alongside, same era.
Towards algorithmic experience: Initial efforts for social media contexts. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . ACM, 286
Oscar Alvarado and Annika Waern. 2018 · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for Deep Neural Networks. In 6th International Conference on Learning Representations (ICLR 2018)
Marco Ancona, Enea Ceolini, Cengiz Oztireli, and Markus Gross. 2018 · 2018
Cited alongside, same era.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, et al · 2018
Cited alongside, same era.
Clinically applicable deep learning for diagnosis and referral in retinal disease
Jeffrey De Fauw, Joseph R Ledsam, Bernardino Romera-Paredes, Stanislav Nikolov, Nenad Tomasev, Sam Blackwell, Harry Askham, Xavier Glorot, Brendan O’Donoghue, Daniel Visentin, et al · 2018
Cited alongside, same era.
Improving simple models with confidence profiles. In Advances in Neural Information Processing Systems . 10296–10306
Amit Dhurandhar, Karthikeyan Shanmugam, Ronny Luss, and Peder A Olsen. 2018 · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA) . IEEE, 80–89
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Cited alongside, same era.
Fair, transparent, and accountable algorithmic decision-making processes
Bruno Lepri, Nuria Oliver, Emmanuel Letouzé, Alex Pentland, and Patrick Vinck. 2018 · 2018
Cited alongside, same era.
Closest in time.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximilian Alber, Christopher J Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel. 2019 · 2019
Closest in time.
On the Connection Between Adversarial Robustness and Saliency Map Interpretability. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.), Vol. 97. PMLR, Long Beach, California, USA, 1823–1832
Christian Etmann, Sebastian Lunz, Peter Maass, and Carola Schoenlieb. 2019 · 2019
Closest in time.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Closest in time.
Interpretable and Differentially Private Predictions
Frederik Harder, Matthias Bauer, and Mijung Park. 2019 · 2019
Closest in time.
Improving fairness in machine learning systems: What do industry practitioners need?. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . ACM, 600
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019 · 2019
Closest in time.
Please Stop Permuting Features: An Explanation and Alternatives
Giles Hooker and Lucas Mentch. 2019 · 2019
Closest in time.
Model Reconstruction from Model Explanations
Smitha Milli, Ludwig Schmidt, Anca Dragan, and Moritz Hardt. 2019 · 2019
Closest in time.
Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency . ACM, 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Closest in time.
Explaining explanations in AI. In Proceedings of the conference on fairness, accountability, and transparency . ACM, 279–288
Brent Mittelstadt, Chris Russell, and Sandra Wachter. 2019 · 2019
Closest in time.
Automatic Model Monitoring for Data Streams
Fábio Pinto, Marco OP Sampaio, and Pedro Bizarro. 2019 · 2019
Closest in time.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Closest in time.
Shubham Sharma, Jette Henderson, and Joydeep Ghosh. 2019 · 2019
Closest in time.
Privacy Risks of Explaining Machine Learning Models
Reza Shokri, Martin Strobel, and Yair Zick. 2019 · 2019
Closest in time.
Understanding Impacts of High-Order Loss Approximations and Features in Deep Learning Interpretation. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.), Vol. 97. PMLR, Long Beach, California, USA, 5848–5856
Sahil Singla, Eric Wallace, Shi Feng, and Soheil Feizi. 2019 · 2019
Closest in time.
Robustness May Be at Odds with Accuracy. In International Conference on Learning Representations
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. 2019 · 2019
Closest in time.
Actionable recourse in linear classification. In Proceedings of the Conference on Fairness, Accountability, and Transparency . ACM, 10–19
Berk Ustun, Alexander Spangher, and Yang Liu. 2019 · 2019
Closest in time.
Transparency: motivations and challenges
Adrian Weller. 2019 · 2019
Closest in time.
The What-If Tool: Interactive Probing of Machine Learning Models
James Wexler, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda Viegas, and Jimbo Wilson. 2019 · 2019
Closest in time.
How Sensitive are Sensitivity-Based Explanations?
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Sai Suggala, David Inouye, and Pradeep Ravikumar. 2019 · 2019
Closest in time.