2021

Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization

Tessera, Kale-ab, Hooker, Sara, Rosman, Benjamin

Understand

Training sparse networks to converge to the same performance as dense neural architectures has proven to be elusive.

  • Recent work suggests that initialization is the key.
  • However, while this direction of research has had some success, focusing on initialization alone appears to be inadequate.
  • In this paper, we take a broader view of training sparse networks and consider the role of regularization, optimization, and architecture choices on sparse models.

Reading the bibliography…