Fetching the paper…
Reading the bibliography…
Most recent work on interpretability of complex machine learning models has focused on estimating $\textit{a posteriori}$ explanations for previously trained models around specific predictions.
“Learning multiple layers of features from tiny images”, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
“Practical Bayesian Optimization of Machine Learning Algorithms”
Jasper Snoek, Hugo Larochelle and Ryan Adams · 2012
Earlier work this paper cites.
“UCI Machine Learning Repository”, 2013
Moshe Lichman and Kevin Bache · 2013
Earlier work this paper cites.
“Deep inside convolutional networks: Visualising image classification models and saliency maps”
Karen Simonyan, Andrea Vedaldi and Andrew Zisserman · 2014
Earlier work this paper cites.
“Visualizing and understanding convolutional networks”
Matthew Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
“On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation”
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus Müller and Wojciech Samek · 2015
Earlier work this paper cites.
“Explaining and Harnessing Adversarial Examples”
Ian. Goodfellow, Jonathon Shlens and Christian Szegedy · 2015
Earlier work this paper cites.
“Understanding neural networks through deep visualization”
Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs and Hod Lipson · 2015
Earlier work this paper cites.
“Rationalizing Neural Predictions”
Tao Lei, Regina Barzilay and Tommi Jaakkola · 2016
Earlier work this paper cites.
“"Why Should I Trust You?": Explaining the Predictions of Any Classifier”
Marco Ribeiro, Sameer Singh and Carlos Guestrin · 2016
Cited alongside, same era.
“Evaluating the visualization of what a deep neural network has learned”
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin and Klaus Müller · 2016
Cited alongside, same era.
“A causal framework for explaining the predictions of black-box sequence-to-sequence models”
David Alvarez-Melis and Tommi. Jaakkola · 2017
Cited alongside, same era.
“"What is relevant in a text document?": An interpretable machine learning approach”
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller and Wojciech Samek · 2017
Cited alongside, same era.
“Towards a Rigorous Science of Interpretable Machine Learning”
Finale Doshi-Velez and Been Kim · 2017
Cited alongside, same era.
“Contextual Explanation Networks”
Maruan Al-Shedivat, Avinava Dubey and Eric Xing · 2017
Later among the works it cites.
“Learning Important Features Through Propagating Activation Differences”
Avanti Shrikumar, Peyton Greenside and Anshul Kundaje · 2017
Later among the works it cites.
“Axiomatic attribution for deep networks”
Mukund Sundararajan, Ankur Taly and Qiqi Yan · 2017
Later among the works it cites.
“From parity to preference-based notions of fairness in classification”
Muhammad Zafar, Isabel Valera, Manuel Rodriguez, Krishna Gummadi and Adrian Weller · 2017
Later among the works it cites.
“On the Robustness of Interpretability Methods”
David Alvarez-Melis and Tommi Jaakkola · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The (Un)reliability of saliency methods”
P.-J. Kindermans, S Hooker, J Adebayo, M Alber, K.\˜T. Schütt, S Dähne, D Erhan and B Kim · 2017
Cited alongside, same era.
“A unified approach to interpreting model predictions”
Scott Lundberg and Su-In Lee · 2017
Cited alongside, same era.
“Methods for interpreting and understanding deep neural networks”
Grégoire Montavon, Wojciech Samek and Klaus-Robert Müller · 2017
Cited alongside, same era.
Ramprasaath. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh and Dhruv Batra · 2017
Cited alongside, same era.
“Beyond Distributive Fairness in Algorithmic Decision Making: Feature Selection for Procedurally Fair Learning”
Nina Grgic-Hlaca, Muhammad Zafar, Krishna Gummadi and Adrian Weller · 2018
Closest in time.
“ Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV) ”
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas and Rory Sayres · 2018
Closest in time.
Oscar Li, Hao Liu, Chaofan Chen and Cynthia Rudin · 2018
Closest in time.