Understand
We propose a novel method for explaining the predictions of any classifier.
- In our approach, local explanations are expected to explain both the outcome of a prediction and how that prediction would change if 'things had been different'.
- Furthermore, we argue that satisfactory explanations cannot be dissociated from a notion and measure of fidelity, as advocated in the early days of neural networks' knowledge extraction.
- We introduce a definition of fidelity to the underlying classifier for local explanation models which is based on distances to a target decision boundary.
Reading the bibliography…