Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Y. Nesterov et al · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer ReLU neural networks
I. Safran and O. Shamir · 2018
Cited alongside, same era.
No spurious local minima in a two hidden unit relu network
C. Wu, J. Luo, and J. D. Lee · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
S. Arora, S. Du, W. Hu, Z. Li, and R. Wang · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Cited alongside, same era.
Towards understanding the importance of shortcut connections in residual networks
T. Liu, M. Chen, M. Zhou, S. S. Du, E. Zhou, and T. Zhao · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. D. Lee, Q. Liu, and T. Ma · 2019
Cited alongside, same era.