Understand
We argue that the optimization plays a crucial role in generalization of deep learning models through implicit regularization.
- We do this by demonstrating that generalization ability is not controlled by network size but rather by some other implicit control.
- We then demonstrate how changing the empirical optimization procedure can improve generalization, even if actual optimization quality is not affected.
- We do so by studying the geometry of the parameter space of deep networks, and devising an optimization algorithm attuned to this geometry.
Built on
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 1999
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
Cryptographic hardness for learning intersections of halfspaces
Adam R Klivansand Alexander A Sherstov · 2006
Earlier work this paper cites.
Similar
Introduction to the Theory of Computation
Michael Sipser · 2006
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Cited alongside, same era.
On the universality of online mirror descent
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2011
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro
Cited in the paper.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro
Cited in the paper.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro
Cited in the paper.
Then
Maxout networks
Ian J. Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron C. Courville, and Yoshua Bengio · 2013
Later among the works it cites.
From average case complexity to improper learning complexity
Amit Daniely, Nati Linial, and Shai Shalev-Shwartz · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…