2018

Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced

Du, Simon S., Hu, Wei, Lee, Jason D.

Understand

We study the implicit regularization imposed by gradient descent for learning multi-layer homogeneous functions including feed-forward fully connected and convolutional deep neural networks with linear, ReLU or Leaky ReLU activation.

  • We rigorously prove that gradient flow (i.e.
  • gradient descent with infinitesimal step size) effectively enforces the differences between squared norms across different layers to remain invariant without any explicit regularization.
  • This result implies that if the weights are initially small, gradient flow automatically balances the magnitudes of all layers.

Reading the bibliography…