2024

Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

Braun, Dan, Taylor, Jordan, Goldowsky-Dill, Nicholas et al.

Understand

Identifying the features learned by neural networks is a core challenge in mechanistic interpretability.

  • Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary that reconstructs a network's internal activations, have been used to identify these features.
  • However, SAEs may learn more about the structure of the datatset than the computational structure of the network.
  • There is therefore only indirect reason to believe that the directions found in these dictionaries are functionally important to the network.

Reading the bibliography…