Fetching the paper…
Reading the bibliography…
Deep neural networks achieve stellar generalisation even when they have enough parameters to easily fit all their training data.
Three unfinished works on the optimal storage capacity of networks
E. Gardner and B. Derrida · 1989
Earlier work this paper cites.
Improving a Network Generalization Ability by Selecting Examples
W. Kinzel, P. Ruján, and P. Rujan · 1990
Earlier work this paper cites.
Statistical mechanics of learning from examples
H. S. Seung, H. Sompolinsky, and N. Tishby · 1992
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
The statistical mechanics of learning a rule
T.L.H. Watkin, A. Rau, and M. Biehl · 1993
Earlier work this paper cites.
Learning by on-line gradient descent
M. Biehl and H. Schwarze · 1995
Earlier work this paper cites.
Exact Solution for On-Line Learning in Multilayer Neural Networks
D. Saad and S.A. Solla · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
D. Saad and S.A. Solla · 1995
Earlier work this paper cites.
On-line backpropagation in two-layered neural networks
P. Riegler and M. Biehl · 1995
Earlier work this paper cites.
Learning with Noise and Regularizers Multilayer Neural Networks
D. Saad and S.A. Solla · 1997
Earlier work this paper cites.
Statistical learning theory
V. Vapnik · 1998
Earlier work this paper cites.
Statistical Mechanics of Learning
A. Engel and C. Van den Broeck · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2003
Earlier work this paper cites.
High-Dimensional Probability
R. Vershynin · 2009
Earlier work this paper cites.
Foundations of Machine Learning
M. Mohri, A. Rostamizadeh, and A. Talwalkar · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A.M. Saxe, James L. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G.E. Hinton · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Norm-Based Capacity Control in Neural Networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
Statistical physics of inference: thresholds and algorithms
L. Zdeborová and F. Krzakala · 2016
Cited alongside, same era.
Statistical mechanics of optimal convex inference in high dimensions
M.S. Advani and S. Ganguli · 2016
Cited alongside, same era.
Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
G.K. Dziugaite and D.M. Roy · 2017
The committee machine: Computational to statistical gaps in learning a two-layers neural network
B. Aubin, A. Maillard, J. Barbier, F. Krzakala, N. Macris, and L. Zdeborová · 2018
Later among the works it cites.
Comparing Dynamics: Deep Neural Networks versus Glassy Systems
M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G.B. Arous, C. Cammarota, Y. LeCun, M. Wyart, and G. Biroli · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P. Nguyen · 2018
Later among the works it cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
G.M. Rotskoff and E. Vanden-Eijnden · 2018
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Cited alongside, same era.
A Closer Look at Memorization in Deep Networks
D. Arpit, S. Jastrz, M.S. Kanwal, T. Maharaj, A. Fischer, A. Courville, and Y. Bengio · 2017
Cited alongside, same era.
Implicit Regularization in Matrix Factorization
S. Gunasekar, B. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Cited alongside, same era.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2017
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
M.S. Advani and A.M. Saxe · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Later among the works it cites.
Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
Y. Li and Y. Liang · 2018
Later among the works it cites.
A Solvable High-Dimensional Model of GAN
C. Wang, Hong Hu, and Yue M. Lu · 2018
Later among the works it cites.
Size-independent sample complexity of neural networks
N. Golowich, A. Rakhlin, and O. Shamir · 2019
Closest in time.
Mean field analysis of neural networks: A central limit theorem
J. Sirignano and K. Spiliopoulos · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
S.S. Du, X. Zhai, B. Poczos, and A. Singh · 2019
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2019
Closest in time.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
A.K. Lampinen and S. Ganguli · 2019
Closest in time.