2019

How noise affects the Hessian spectrum in overparameterized neural networks

Wei, Mingwei, Schwab, David J

Understand

Stochastic gradient descent (SGD) forms the core optimization method for deep neural networks.

  • While some theoretical progress has been made, it still remains unclear why SGD leads the learning dynamics in overparameterized networks to solutions that generalize well.
  • Here we show that for overparameterized networks with a degenerate valley in their loss landscape, SGD on average decreases the trace of the Hessian of the loss.
  • We also generalize this result to other noise structures and show that isotropic noise in the non-degenerate subspace of the Hessian decreases its determinant.

Reading the bibliography…