Fetching the paper…
Reading the bibliography…
A main puzzle of deep neural networks (DNNs) revolves around the apparent absence of "overfitting", defined in this paper as follows: the expected error does not get worse when increasing the number of neurons or of iterations of gradient descent.
Differential Equations with Discontinuous Righthand Sides: Control Systems
F.M. Arscott and A.F. Filippov · 1988
Earlier work this paper cites.
The hartman-grobman theorem for caratheodory-type differential equations in banach spaces
Thomas Wanner · 2000
Earlier work this paper cites.
Neural Network Learning - Theoretical Foundations
M. Anthony and P. Bartlett · 2002
Earlier work this paper cites.
Stability Theory of Dynamical Systems
N.P. Bhatia and G.P. Szegö · 2002
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Theory I: Why and when can deep - but not shallow - networks avoid the curse of dimensionality
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao · 2016
Cited alongside, same era.
Deep vs. shallow networks: An approximation theory perspective
H.N. Mhaskar and T. Poggio · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and P. Kingma Diederik · 2016
Cited alongside, same era.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
The Implicit Bias of Gradient Descent on Separable Data
D. Soudry, E. Hoffer, and N. Srebro · 2017
Cited alongside, same era.
Musings on deep learning: Optimization properties of SGD
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N.Golowich, and T. Poggio · 2017
Later among the works it cites.
Waiting for godot
L. Rosasco and B. Recht · 2017
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2017
Later among the works it cites.
Theory of deep learning III: explaining the non-overfitting puzzle
T. Poggio, Q. Liao, B. Miranda, L. Rosasco, X. Boix, J. Hidary, and H. Mhaskar · 2017
Later among the works it cites.
Theory of deep learning IIb: Optimization properties of SGD
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N.Golowich, and T. Poggio · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel Rodrigues · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Cited alongside, same era.
T. Poggio and Q. Liao · 2017
Later among the works it cites.
To understand deep learning we need to understand kernel learning
M. Belkin, S. Ma, and S. Mandal · 2018
Closest in time.
Implicit Bias of Gradient Descent on Linear Convolutional Networks
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Closest in time.