2023

The LLM Surgeon

van der Ouderaa, Tycho F. A., Nagel, Markus, van Baalen, Mart et al.

Understand

State-of-the-art language models are becoming increasingly large in an effort to achieve the highest performance on large corpora of available textual data.

  • However, the sheer size of the Transformer architectures makes it difficult to deploy models within computational, environmental or device-specific constraints.
  • We explore data-driven compression of existing pretrained models as an alternative to training smaller models from scratch.
  • To do so, we scale Kronecker-factored curvature approximations of the target loss landscape to large language models.

Reading the bibliography…