2017

A Regularized Framework for Sparse and Structured Neural Attention

Niculae, Vlad, Blondel, Mathieu

Understand

Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input.

  • We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator.
  • We show that the gradient of this operator defines a mapping from real values to probabilities, suitable as an attention mechanism.
  • Our framework includes softmax and a slight generalization of the recently-proposed sparsemax as special cases.

Reading the bibliography…