2022

Impossibility Theorems for Feature Attribution

Bilodeau, Blair, Jaques, Natasha, Koh, Pang Wei et al.

Understand

Despite a sea of interpretability methods that can produce plausible explanations, the field has also empirically seen many failure cases of such methods.

  • In light of these results, it remains unclear for practitioners how to use these methods and choose between them in a principled way.
  • In this paper, we show that for moderately rich model classes (easily satisfied by neural networks), any feature attribution method that is complete and linear -- for example, Integrated Gradients and SHAP -- can provably fail to improve on random guessing for inferring model behaviour.
  • Our results apply to common end-tasks such as characterizing local model behaviour, identifying spurious features, and algorithmic recourse.

Reading the bibliography…