2022

Token Dropping for Efficient BERT Pretraining

Hou, Le, Pang, Richard Yuanzhe, Zhou, Tianyi et al.

Understand

Transformer-based models generally allocate the same amount of computation for each token in a given sequence.

  • We develop a simple but effective "token dropping" method to accelerate the pretraining of transformer models, such as BERT, without degrading its performance on downstream tasks.
  • In short, we drop unimportant tokens starting from an intermediate layer in the model to make the model focus on important tokens; the dropped tokens are later picked up by the last layer of the model so that the model still produces full-length sequences.
  • We leverage the already built-in masked language modeling (MLM) loss to identify unimportant tokens with practically no computational overhead.

Reading the bibliography…