2017

High-dimensional dynamics of generalization error in neural networks

Advani, Madhu S., Saxe, Andrew M.

Understand

We perform an average case analysis of the generalization dynamics of large neural networks trained using gradient descent.

  • We study the practically-relevant "high-dimensional" regime where the number of free parameters in the network is on the order of or even larger than the number of examples in the dataset.
  • Using random matrix theory and exact solutions in linear models, we derive the generalization error and training error dynamics of learning and analyze how they depend on the dimensionality of data and signal to noise ratio of the learning problem.
  • We find that the dynamics of gradient descent learning naturally protect against overtraining and overfitting in large networks.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…