Fetching the paper…
Reading the bibliography…
It has been observed \citep{zhang2016understanding} that deep neural networks can memorize: they achieve 100\% accuracy on training data.
On exact computation with an infinitely wide neural net
S. Arora, S. S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang · 1904
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 1905
Earlier work this paper cites.
Distributional and l q l^{q} norm inequalities for polynomials over convex bodies in r n r^{n}
A. Carbery and J. Wright · 2001
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
D. A. Spielman and S.-H. Teng · 2004
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
M. Rudelson and R. Vershynin · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
J. Sun, Q. Qu, and J. Wright · 2015
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
R. Ge, J. D. Lee, and T. Ma · 2016
Earlier work this paper cites.
Polynomial-time tensor decompositions with sum-of-squares
T. Ma, J. Shi, and D. Steurer · 2016
Earlier work this paper cites.
Non-square matrix sensing without spurious local minima via the burer-monteiro approach
D. Park, A. Kyrillidis, C. Caramanis, and S. Sanghavi · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
N. Agarwal, Z. Allen-Zhu, B. Bullins, E. Hazan, and T. Ma · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
R. Ge, C. Jin, and Y. Zheng · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2018
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Later among the works it cites.
Small relu networks are powerful memorizers: a tight analysis of memorization capacity
C. Yun, S. Sra, and A. Jadbabaie · 2018
Later among the works it cites.
Does data interpolation contradict statistical optimality?
M. Belkin, A. Rakhlin, and A. B. Tsybakov · 2019
Closest in time.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Belkin, D. J. Hsu, and P. Mitra · 2018
Cited alongside, same era.
Accelerated methods for nonconvex optimization
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
S. S. Du and J. D. Lee · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang
Cited in the paper.
Closest in time.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
T. Liang, A. Rakhlin, and X. Zhai · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Closest in time.
S. Oymak and M. Soltanolkotabi · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
D. Zou and Q. Gu · 2019
Closest in time.