2017

Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Kim, Been, Wattenberg, Martin, Gilmer, Justin et al.

Understand

The interpretation of deep learning models is a challenge due to their size, complexity, and often opaque internal state.

  • In addition, many systems, such as image classifiers, operate on low-level features rather than high-level concepts.
  • To address these challenges, we introduce Concept Activation Vectors (CAVs), which provide an interpretation of a neural net's internal state in terms of human-friendly concepts.
  • The key idea is to view the high-dimensional internal state of a neural net as an aid, not an obstacle.

Reading the bibliography…