2022

Interpreting Neural Networks through the Polytope Lens

Black, Sid, Sharkey, Lee, Grinsztajn, Leo et al.

Understand

Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level.

  • What are the fundamental primitives of neural network representations? Previous mechanistic descriptions have used individual neurons or their linear combinations to understand the representations a network has learned.
  • But there are clues that neurons and their linear combinations are not the correct fundamental units of description: directions cannot describe how neural networks use nonlinearities to structure their representations.
  • Moreover, many instances of individual neurons and their combinations are polysemantic (i.e.

Reading the bibliography…