2018

Minimum weight norm models do not always generalize well for over-parameterized problems

Shah, Vatsal, Kyrillidis, Anastasios, Sanghavi, Sujay

Understand

This work is substituted by the paper in arXiv:2011.14066.

  • Stochastic gradient descent is the de facto algorithm for training deep neural networks (DNNs).
  • Despite its popularity, it still requires fine tuning in order to achieve its best performance.
  • This has led to the development of adaptive methods, that claim automatic hyper-parameter optimization.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…