Fetching the paper…
Reading the bibliography…
The key to generalization is controlling the complexity of the network.
Differential Equations with Discontinuous Righthand Sides: Control Systems
F.M. Arscott and A.F. Filippov · 1988
Earlier work this paper cites.
On the relationship between generalization error, hypothesis complexity, and sample complexity for radial basis functions
P. Niyogi and F. Girosi · 1996
Earlier work this paper cites.
The existence and uniqueness of the minimum norm solution to certain linear and nonlinear problems
Paulo Jorge S. G. Ferreira · 1996
Earlier work this paper cites.
On gradient adaptation with unit-norm constraints
S. C. Douglas, S. Amari, and S. Y. Kung · 2000
Earlier work this paper cites.
The mathematics of learning: Dealing with data
T. Poggio and S. Smale · 2003
Earlier work this paper cites.
Margin maximizing loss functions
Saharon Rosset, Ji Zhu, and Trevor Hastie · 2003
Earlier work this paper cites.
Introduction to statistical learning theory
O. Bousquet, S. Boucheron, and G. Lugosi · 2003
Earlier work this paper cites.
Convexity, classification, and risk bounds
Peter L. Bartlett, Michael I. Jordan, and Jon D. McAuliffe · 2003
Earlier work this paper cites.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Learning with incremental iterative regularization
Lorenzo Rosasco and Silvia Villa · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingm · 2016
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Earlier work this paper cites.
Robust large margin deep neural networks
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel Rodrigues · 2017
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Earlier work this paper cites.
Musings on deep learning: Optimization properties of SGD
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N.Golowich, and T. Poggio · 2017
Earlier work this paper cites.
The Implicit Bias of Gradient Descent on Separable Data
D. Soudry, E. Hoffer, and N. Srebro · 2017
Cited alongside, same era.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason D. Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
How Does Batch Normalization Help Optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
Just Interpolate: Kernel ”Ridgeless” Regression Can Generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Later among the works it cites.
Consistency of Interpolation with Laplace Kernels is a High-Dimensional Phenomenon
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L. Bartlett, and Inderjit S. Dhillon · 2017
Cited alongside, same era.
Learning non-overlapping convolutional neural networks with multiple kernels
Kai Zhong, Zhao Song, and Inderjit S. Dhillon · 2017
Cited alongside, same era.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Theory of deep learning IIb: Optimization properties of SGD
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N.Golowich, and T. Poggio · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient langevin dynamics: A nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Cited alongside, same era.
Theory II: Landscape of the empirical risk in deep learning
T. Poggio and Q. Liao · 2017
Cited alongside, same era.
Alexander Rakhlin and Xiyu Zhai · 2018
Later among the works it cites.
A Mean Field View of the Landscape of Two-Layers Neural Networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Later among the works it cites.
A surprising linear relationship predicts test performance in deep networks
Qianli Liao, Brando Miranda, Andrzej Banburski, Jack Hidary, and Tomaso A. Poggio · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhi yuan Li, and Ruosong Wang · 2019
Closest in time.
Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
Mor Shpigel Nacson, Suriya Gunasekar, Jason D. Lee, Nathan Srebro, and Daniel Soudry · 2019
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Closest in time.
Theory of deep learning III: Dynamics and generalization in deep networks
A. Banburski, Q. Liao, B. Miranda, T. Poggio, L. Rosasco, B. Liang, and J. Hidary · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Closest in time.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Closest in time.
Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
Phan-Minh Nguyen · 2019
Closest in time.