Fetching the paper…
Reading the bibliography…
While over-parameterization is widely believed to be crucial for the success of optimization for the neural networks, most existing theories on over-parameterization do not fully explain the reason -- they either work in the Neural Tangent Kernel regime where neurons don't move much, or require an enormous number of neurons.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R. (2019b) · 1901
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R. (2019a) · 1904
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q. (2019) · 1905
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Araújo, D., Oliveira, R. I., and Yukimura, D. (2019) · 1906
Earlier work this paper cites.
Sparse optimization on measures with over-parameterized gradient descent
Chizat, L. (2019) · 1907
Earlier work this paper cites.
Student specialization in deep relu networks with finite width and input dimension
Tian, Y. (2019) · 1909
Earlier work this paper cites.
Handbook of mathematical functions with formulas, graphs, and mathematical tables
Abramowitz, M. and Stegun, I. A. (1948) · 1948
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels
Lojasiewicz, S. (1963) · 1963
Earlier work this paper cites.
A rigorous framework for the mean field limit of multilayer neural networks
Nguyen, P.-M. and Pham, H. T. (2020) · 2001
Earlier work this paper cites.
Safran, I., Yehudai, G., and Shamir, O. (2020) · 2006
Earlier work this paper cites.
Modeling from features: a mean-field framework for over-parameterized deep neural networks
Fang, C., Lee, J. D., Yang, P., and Zhang, T. (2020) · 2007
Earlier work this paper cites.
Learning over-parametrized two-layer relu neural networks beyond ntk
Li, Y., Ma, T., and Zhang, H. R. (2020) · 2007
Earlier work this paper cites.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A. (2012) · 2012
Earlier work this paper cites.
Analysis of boolean functions
O’Donnell, R. (2014) · 2014
Cited alongside, same era.
Daniely, A., Frostig, R., and Singer, Y. (2016) · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M. (2016) · 2016
Cited alongside, same era.
Stochastic particle gradient descent for infinite ensembles
Nitanda, A. and Suzuki, T. (2017) · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S. (2017) · 2017
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M. (2018) · 2018
Later among the works it cites.
Trainability and accuracy of neural networks: An interacting particle system approach
Rotskoff, G. M. and Vanden-Eijnden, E. (2018) · 2018
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F. (2019) · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X. (2019) · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Mei, S., Misiakiewicz, T., and Montanari, A. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z. (2018) · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F. (2018) · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2018) · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M. (2018) · 2018
Cited alongside, same era.
Learning two-layer neural networks with symmetric inputs
Ge, R., Kuditipudi, R., Li, Z., and Wang, X. (2018) · 2018
Cited alongside, same era.
Learning one convolutional layer with overlapping patches
Goel, S., Klivans, A., and Meka, R. (2018) · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Cited alongside, same era.
Wei, C., Lee, J. D., Liu, Q., and Ma, T. (2019) · 2019
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Zhang, X., Yu, Y., Wang, L., and Gu, Q. (2019) · 2019
Later among the works it cites.
Algorithms and sq lower bounds for pac learning one-hidden-layer relu networks
Diakonikolas, I., Kane, D. M., Kontonis, V., and Zarifis, N. (2020) · 2020
Later among the works it cites.
A mean field analysis of deep resnet and beyond: Towards provably optimization via overparameterization from depth
Lu, Y., Ma, C., Lu, Y., Lu, J., and Ying, L. (2020) · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Oymak, S. and Soltanolkotabi, M. (2020) · 2020
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem
Sirignano, J. and Spiliopoulos, K. (2020) · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q. (2020) · 2020
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A. (2021) · 2021
Closest in time.