2020

What Happens To BERT Embeddings During Fine-tuning?

Merchant, Amil, Rahimtoroghi, Elahe, Pavlick, Ellie et al.

Understand

While there has been much recent work studying how linguistic information is encoded in pre-trained sentence representations, comparatively little is understood about how these models change when adapted to solve downstream tasks.

  • Using a suite of analysis techniques (probing classifiers, Representational Similarity Analysis, and model ablations), we investigate how fine-tuning affects the representations of the BERT model.
  • We find that while fine-tuning necessarily makes significant changes, it does not lead to catastrophic forgetting of linguistic phenomena.
  • We instead find that fine-tuning primarily affects the top layers of BERT, but with noteworthy variation across tasks.

Reading the bibliography…