2019

Adaptively Sparse Transformers

Correia, Gonçalo M., Niculae, Vlad, Martins, André F. T.

Understand

Attention mechanisms have become ubiquitous in NLP.

  • Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed attention.
  • The multiple heads learn diverse types of word relationships.
  • However, with standard softmax attention, all attention heads are dense, assigning a non-zero weight to all context words.

Reading the bibliography…