Gradient descent provably optimizes over-parameterized neural networks
Original
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Original
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layers neural networks
Mei, S., Montanari, A., and Nguyen, P · 2018
Cited alongside, same era.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Original
Rotskoff, G. M. and Vanden-Eijnden, E · 2018
Cited alongside, same era.
Mean field analysis of neural networks
Original
Sirignano, J. and Spiliopoulos, K · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Allen-Zhu, Z. and Li, Y · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.