Fetching the paper…
Reading the bibliography…
Deep nets generalize well despite having more parameters than the number of training samples.
Relating data compression and learnability
Nick Littlestone and Manfred Warmuth · 1986
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Some PAC-Bayesian theorems
David A McAllester · 1998
Earlier work this paper cites.
The connection between regularization operators and support vector kernels
Alex J Smola, Bernhard Schölkopf, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Algorithmic stability and sanity-check bounds for leave-one-out cross-validation
Michael Kearns and Dana Ron · 1999
Earlier work this paper cites.
PAC-Bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Regularization networks and support vector machines
Theodoros Evgeniou, Massimiliano Pontil, and Tomaso Poggio · 2000
Earlier work this paper cites.
A rank minimization heuristic with application to minimum order system approximation
Maryam Fazel, Haitham Hindi, and Stephen P Boyd · 2001
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Maximum-margin matrix factorization
Nathan Srebro, Jason Rennie, and Tommi S Jaakkola · 2005
Earlier work this paper cites.
Random projection, margins, kernels, and feature-selection
Avrim Blum · 2006
Cited alongside, same era.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Cited alongside, same era.
Universal donsker classes and metric entropy
Richard M Dudley · 2010
Cited alongside, same era.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Later among the works it cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J Foster, and Matus Telgarsky · 2017
Later among the works it cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Hoeffding’s inequality for sums of weakly dependent random variables
Christos Pelekis and Jan Ramon · 2015
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, and Yann LeCun · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2016
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro
Cited in the paper.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Later among the works it cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Later among the works it cites.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2018
Closest in time.
On the importance of single directions for generalization
Ari Morcos, David GT Barrett, Matthew Botvinick, and Neil Rabinowitz · 2018
Closest in time.