2019

Measurable Counterfactual Local Explanations for Any Classifier

White, Adam, Garcez, Artur d'Avila

Understand

We propose a novel method for explaining the predictions of any classifier.

  • In our approach, local explanations are expected to explain both the outcome of a prediction and how that prediction would change if 'things had been different'.
  • Furthermore, we argue that satisfactory explanations cannot be dissociated from a notion and measure of fidelity, as advocated in the early days of neural networks' knowledge extraction.
  • We introduce a definition of fidelity to the underlying classifier for local explanation models which is based on distances to a target decision boundary.

Reading the bibliography…