2019

Why gradient clipping accelerates training: A theoretical justification for adaptivity

Zhang, Jingzhao, He, Tianxing, Sra, Suvrit et al.

Understand

We provide a theoretical explanation for the effectiveness of gradient clipping in training deep neural networks.

  • The key ingredient is a new smoothness condition derived from practical neural network training examples.
  • We observe that gradient smoothness, a concept central to the analysis of first-order optimization algorithms that is often assumed to be a constant, demonstrates significant variability along the training trajectory of deep neural networks.
  • Further, this smoothness positively correlates with the gradient norm, and contrary to standard assumptions in the literature, it can grow with the norm of the gradient.

Reading the bibliography…