Fetching the paper…
Reading the bibliography…
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex.
Gaussian sobolev spaces and stochastic calculus of variations
Malliavin, P · 1995
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
Learning polynomials with neural networks
Andoni, A., Panigrahy, R., Valiant, G., and Zhang, L · 2014
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, B. D. and Vidal, R · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2016
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Earlier work this paper cites.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Kleinberg, J., Ganguli, S., and Sohl-Dickstein, J · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Safran, I. and Shamir, O · 2016
Earlier work this paper cites.
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Globally optimal gradient descent for a ConvNet with gaussian inputs
Brutzkus, A. and Globerson, A · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2017
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Closest in time.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z · 2018
Closest in time.
A mean field view of the landscape of two-layers neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Closest in time.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Mou, W., Wang, L., Zhai, X., and Zheng, K · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Li, Y. and Yuan, Y · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Cited alongside, same era.
Learning ReLUs via gradient descent
Soltanolkotabi, M · 2017
Cited alongside, same era.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Tian, Y · 2017
Cited alongside, same era.
Rotskoff, G. M. and Vanden-Eijnden, E · 2018
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
Safran, I. and Shamir, O · 2018
Closest in time.
Mean field analysis of neural networks
Sirignano, J. and Spiliopoulos, K · 2018
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2018
Closest in time.
Neural networks with finite intrinsic dimension have no spurious valleys
Venturi, L., Bandeira, A., and Bruna, J · 2018
Closest in time.
On the margin theory of feedforward neural networks
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2018
Closest in time.
Learning one-hidden-layer relu networks via gradient descent
Zhang, X., Yu, Y., Wang, L., and Gu, Q · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Closest in time.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 2019
Closest in time.
Residual learning without normalization via better initialization
Zhang, H., Dauphin, Y. N., and Ma, T · 2019
Closest in time.