Understand
Large-scale pretraining of transformers has been central to the success of foundation models.
- However, the scale of those models limits our understanding of the mechanisms at play during optimization.
- In this work, we study the training dynamics of transformers in a controlled and interpretable setting.
- On the sparse modular addition task, we demonstrate that specialized attention circuits, called clustering heads, can be implemented during gradient descent to solve the problem.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…