Fetching the paper…
Reading the bibliography…
Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the Neural Tangent Kernel (NTK).
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
On Exact Computation with an Infinitely Wide Neural Net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
An elementary introduction to modern convex geometry
Keith Ball · 1997
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii, Dmitry Panchenko, et al · 2002
Earlier work this paper cites.
Feature selection, l 1 vs. l 2 regularization, and rotational invariance
Andrew Y Ng · 2004
Earlier work this paper cites.
1-norm support vector machines
Ji Zhu, Saharon Rosset, Robert Tibshirani, and Trevor J Hastie · 2004
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2006
Earlier work this paper cites.
l1 regularization in infinite dimensional feature spaces
Saharon Rosset, Grzegorz Swirszcz, Nathan Srebro, and Ji Zhu · 2007
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson, Roman Vershynin, et al · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Closest in time.
Just Interpolate: Kernel “Ridgeless” Regression Can Generalize
T. Liang and A. Rakhlin · 2018
Closest in time.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Closest in time.
A mean field view of the landscape of two-layers neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Closest in time.
On the importance of single directions for generalization
Ari S Morcos, David GT Barrett, Neil C Rabinowitz, and Matthew Botvinick · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
Simon S Du, Jason D Lee, Yuandong Tian, Barnabas Poczos, and Aarti Singh · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Cong Ma, Kaizheng Wang, Yuejie Chi, and Yuxin Chen · 2017
Cited alongside, same era.
Closest in time.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry · 2018
Closest in time.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Closest in time.
Bayesian convolutional neural networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Closest in time.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Closest in time.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Closest in time.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2018
Closest in time.
Neural networks with finite intrinsic dimension have no spurious valleys
Luca Venturi, Afonso Bandeira, and Joan Bruna · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Closest in time.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.
Xialiang Dou and Tengyuan Liang · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Closest in time.
Data-dependent Sample Complexity of Deep Neural Networks via Lipschitz Augmentation
Colin Wei and Tengyu Ma · 2019
Closest in time.
Greg Yang · 2019
Closest in time.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Closest in time.