Fetching the paper…
Reading the bibliography…
The evolution of a deep neural network trained by the gradient descent can be described by its neural tangent kernel (NTK) as introduced in [20], where it was proven that in the infinite width limit the NTK converges to an explicit limiting kernel and it stays constant during training.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, B. Kingsbury, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2012
Earlier work this paper cites.
Deep convolutional neural networks for lvcsr
T. N. Sainath, A.-r. Mohamed, B. Kingsbury, and B. Ramabhadran · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Ben Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
R. Ge, J. D. Lee, and T. Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Earlier work this paper cites.
Non-square matrix sensing without spurious local minima via the burer-monteiro approach
D. Park, A. Kyrillidis, C. Caramanis, and S. Sanghavi · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Complete dictionary recovery over the sphere i: Overview and the geometric picture
J. Sun, Q. Qu, and J. Wright · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Cited alongside, same era.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2018
Later among the works it cites.
What can resnet learn efficiently, going beyond kernels?
Z. Allen-Zhu and Y. Li · 2019
Closest in time.
A mean-field limit for certain deep neural networks
D. Araújo, R. I. Oliveira, and D. Yukimura · 2019
Closest in time.
On exact computation with an infinitely wide neural net
S. Arora, S. S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Allen-Zhu, Y. Li, and Y. Liang · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 2019
Closest in time.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
K. Kawaguchi and J. Huang · 2019
Closest in time.
Elimination of all bad local minima in deep learning
K. Kawaguchi and L. P. Kaelbling · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. S. Schoenholz, Y. Bahri, J. Sohl-Dickstein, and J. Pennington · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
Mean field limit of the learning dynamics of multilayer neural networks
P.-M. Nguyen · 2019
Closest in time.
Mean field analysis of deep neural networks
J. Sirignano and K. Spiliopoulos · 2019
Closest in time.
Quadratic suffices for over-parametrization via matrix chernoff bound
Z. Song and X. Yang · 2019
Closest in time.
G. Yang · 2019
Closest in time.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 2019
Closest in time.