2024

A Mechanistic Study of Transformers Training Dynamics

Odonnat, Ambroise, Bouaziz, Wassim, Cabannes, Vivien

Understand

Large-scale pretraining of transformers has been central to the success of foundation models.

  • However, the scale of those models limits our understanding of the mechanisms at play during optimization.
  • In this work, we study the training dynamics of transformers in a controlled and interpretable setting.
  • On the sparse modular addition task, we demonstrate that specialized attention circuits, called clustering heads, can be implemented during gradient descent to solve the problem.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…