2022

Oolong: Investigating What Makes Transfer Learning Hard with Controlled Studies

Wu, Zhengxuan, Tamkin, Alex, Papadimitriou, Isabel

Understand

When we transfer a pretrained language model to a new language, there are many axes of variation that change at once.

  • To disentangle the impact of different factors like syntactic similarity and vocabulary similarity, we propose a set of controlled transfer studies: we systematically transform the language of the GLUE benchmark, altering one axis of crosslingual variation at a time, and then measure the resulting drops in a pretrained model's downstream performance.
  • We find that models can largely recover from syntactic-style shifts, but cannot recover from vocabulary misalignment and embedding matrix re-initialization, even with continued pretraining on 15 million tokens.
  • %On the other hand, transferring to a dataset with an unaligned vocabulary is extremely hard to recover from in the low-data regime.

Reading the bibliography…