Fetching the paper…
Reading the bibliography…
One of the defining properties of deep learning is that models are chosen to have many more parameters than available training data.
“(Not) Bounding the True Error”
John Langford and Rich Caruana · 1968
Earlier work this paper cites.
“A Universal Prior for Integers and Estimation by Minimum Description Length”
Jorma Rissanen · 1983
Earlier work this paper cites.
“Visual Reconstruction”
Andrew Blake and Andrew Zisserman · 1987
Earlier work this paper cites.
“Keeping the Neural Networks Simple by Minimizing the Description Length of the Weights”
Geoffrey. Hinton and Drew van Camp · 1993
Earlier work this paper cites.
“For valid generalization the size of the weights is more important than the size of the network”
Peter Bartlett · 1997
Earlier work this paper cites.
“Flat Minima”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network”
Peter Bartlett · 1998
Earlier work this paper cites.
“PAC-Bayesian Model Averaging”
David. McAllester · 1999
Earlier work this paper cites.
“Bounds for Averaging Classifiers”, 2001
John Langford and Matthias Seeger · 2001
Earlier work this paper cites.
“Rademacher and Gaussian complexities: Risk bounds and structural results”
Peter Bartlett and Shahar Mendelson · 2002
Cited alongside, same era.
“Empirical margin distributions and bounding the generalization error of combined classifiers”
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Cited alongside, same era.
“Quantitatively tight sample complexity bounds”, 2002
John Langford · 2002
Cited alongside, same era.
“The minimum description length principle”
Peter Grünwald · 2007
Cited alongside, same era.
“MNIST handwritten digit database”, http://yann.lecun.com/exdb/mnist/, 2010
Yann LeCun, Corinna Cortes and Christopher.. Burges · 2010
Cited alongside, same era.
“Norm-based capacity control in neural networks”
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2015
Later among the works it cites.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems” Software available from tensorflow.org, 2015
Martin Abadi et al · 2015
Later among the works it cites.
“Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes”
Carlo Baldassi et al · 2016
Later among the works it cites.
“On Graduated Optimization for Stochastic Non-Convex Problems”
Elad Hazan, Kfir Levy and Shai Shalev-Shwartz · 2016
Later among the works it cites.
“The impact of the nonlinearity on the VC-dimension of a deep network”, Preprint, 2017
Peter. Bartlett · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2014
Cited alongside, same era.
“Subdominant Dense Clusters Allow for Simple Learning and High Computational Performance in Neural Networks with Discrete Synapses”
Carlo Baldassi et al · 2015
Cited alongside, same era.
“Train faster, generalize better: Stability of stochastic gradient descent”
Moritz Hardt, Benjamin Recht and Yoram Singer · 2015
Cited alongside, same era.
“Path-SGD: Path-Normalized Optimization in Deep Neural Networks”
Behnam Neyshabur, Ruslan Salakhutdinov and Nathan Srebro · 2015
Cited alongside, same era.
Pratik Chaudhari et al · 2017
Closest in time.
“On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima”
Nitish Keskar et al · 2017
Closest in time.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang et al · 2017
Closest in time.