Fetching the paper…
Reading the bibliography…
We give a simple proof for the global convergence of gradient descent in training deep ReLU networks with the standard square loss, and show some of its improvements over the state-of-the-art.
Quadratic suffices for over-parametrization via matrix chernoff bound, 2020
Song, Z. and Yang, X · 1906
Earlier work this paper cites.
How much over-parameterization is sufficient to learn deep relu networks?, 2019
Chen, Z., Cao, Y., Zou, D., and Gu, Q · 1911
Earlier work this paper cites.
Gradient methods for minimizing functionals
Polyak, B. T · 1963
Earlier work this paper cites.
Local operator theory, random matrices and banach spaces
Davidson, K. R. and Szarek, S. J · 2001
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Daniely, A · 2017
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Vershynin, R · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X
An improved analysis of training over-parameterized deep neural networks
Zou, D. and Gu, Q · 2019
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Huang, J. and Yau, H.-T · 2020
Later among the works it cites.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Nguyen, Q. and Mondelli, M · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Oymak, S. and Soltanolkotabi, M · 2020
Later among the works it cites.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks
Nguyen, Q., Mondelli, M., and Montufar, G · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited in the paper.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A
Cited in the paper.