2023

Successor Heads: Recurring, Interpretable Attention Heads In The Wild

Gould, Rhys, Ong, Euan, Ogden, George et al.

Understand

In this work we present successor heads: attention heads that increment tokens with a natural ordering, such as numbers, months, and days.

  • For example, successor heads increment 'Monday' into 'Tuesday'.
  • We explain the successor head behavior with an approach rooted in mechanistic interpretability, the field that aims to explain how models complete tasks in human-understandable terms.
  • Existing research in this area has found interpretable language model components in small toy models.

Reading the bibliography…