Fetching the paper…
Reading the bibliography…
Many modern neural network architectures are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset.
Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen
J. Schur · 1911
Earlier work this paper cites.
Generalized gradients and applications
Frank H. Clarke · 1975
Earlier work this paper cites.
The concentration of measure phenomenon
M. Ledoux · 2001
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
The landscape of empirical risk for non-convex losses
Song Mei, Yu Bai, and Andrea Montanari · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
The landscape of deep learning algorithms
Pan Zhou and Jiashi Feng · 2017
Earlier work this paper cites.
Learning ReLUs via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with Gaussian inputs
A. Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2017
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J. Foster, and Matus Telgarsky · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Cited alongside, same era.
Spurious valleys in two-layer neural network optimization landscapes
L. Venturi, A. Bandeira, and J. Bruna · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Mei Song, A Montanari, and P Nguyen · 2018
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Cited alongside, same era.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Cited alongside, same era.
The global optimization geometry of shallow linear neural networks
Zhihui Zhu, Daniel Soudry, Yonina C. Eldar, and Michael B. Wakin · 2018
Later among the works it cites.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Later among the works it cites.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Later among the works it cites.
Does data interpolation contradict statistical optimality?
Mikhail Belkin, Alexander Rakhlin, and Alexandre B. Tsybakov · 2018
Later among the works it cites.
Just interpolate: Kernel "ridgeless" regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Grant M. Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Later among the works it cites.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Mikhail Belkin, Daniel Hsu, and Partha Mitra · 2018
Later among the works it cites.