2023

Generating Interpretable Networks using Hypernetworks

Liao, Isaac, Liu, Ziming, Tegmark, Max

Understand

An essential goal in mechanistic interpretability to decode a network, i.e., to convert a neural network's raw weights to an interpretable algorithm.

  • Given the difficulty of the decoding problem, progress has been made to understand the easier encoding problem, i.e., to convert an interpretable algorithm into network weights.
  • Previous works focus on encoding existing algorithms into networks, which are interpretable by definition.
  • However, focusing on encoding limits the possibility of discovering new algorithms that humans have never stumbled upon, but that are nevertheless interpretable.

Reading the bibliography…