Fetching the paper…
Reading the bibliography…
Neural networks are among the most accurate supervised learning methods in use today, but their opacity makes them difficult to trust in critical applications, especially when conditions in training differ from those in test.
Improving generalization performance using double backpropagation
Harris Drucker and Yann Le Cun · 1992
Earlier work this paper cites.
Explanation and understanding
Frank C Keil · 2006
Earlier work this paper cites.
Using ”annotator rationales” to improve machine learning for text categorization
Omar Zaidan, Jason Eisner, and Christine D Piatko · 2007
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert MÞller · 2010
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun, Corinna Cortes, and Christopher J.C. Burges · 2010
Earlier work this paper cites.
Annotator rationales for visual recognition
Jeff Donahue and Kristen Grauman · 2011
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel · 2012
Earlier work this paper cites.
UCI machine learning repository
M. Lichman · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Cited alongside, same era.
Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad · 2015
Cited alongside, same era.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky · 2015
Cited alongside, same era.
Auditing black-box models by obscuring features
Philip Adler, Casey Falk, Sorelle A Friedler, Gabriel Rybeck, Carlos Scheidegger, Brandon Smith, and Suresh Venkatasubramanian · 2016
Cited alongside, same era.
Interpretation of prediction models using the input gradient
Grad-CAM: Why did you say that?
Ramprasaath R Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra · 2016
Later among the works it cites.
Not just a black box: Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje · 2016
Later among the works it cites.
Programs as black-box explanations
Sameer Singh, Marco Tulio Ribeiro, and Carlos Guestrin · 2016
Later among the works it cites.
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yotam Hechtlinger · 2016
Cited alongside, same era.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola · 2016
Cited alongside, same era.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky · 2016
Cited alongside, same era.
The mythos of model interpretability
Zachary C Lipton · 2016
Cited alongside, same era.
An unexpected unity among methods for interpreting model predictions
Scott Lundberg and Su-In Lee · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Cited alongside, same era.
Why should I trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rodriguez, and Krishna P Gummadi · 2016
Later among the works it cites.
Rationale-augmented convolutional neural networks for text classification
Ye Zhang, Iain Marshall, and Byron C Wallace · 2016
Later among the works it cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Closest in time.
Interpretable explanations of black boxes by meaningful perturbation
Ruth Fong and Andrea Vedaldi · 2017
Closest in time.
Autograd
Dougal Mclaurin, David Duvenaud, and Matt Johnson · 2017
Closest in time.
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller · 2017
Closest in time.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Closest in time.