2023

A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations

Chughtai, Bilal, Chan, Lawrence, Nanda, Neel

Understand

Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks.

  • In this work, we study the universality hypothesis by examining how small neural networks learn to implement group composition.
  • We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory.
  • We then show that networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations.

Reading the bibliography…