Fetching the paper…
Reading the bibliography…
In this paper, we analyze the landscape of the true loss of neural networks with one hidden layer and ReLU, leaky ReLU, or quadratic activation.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P., and Hornik, K · 1989
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Fukumizu, K., and Amari, S.-i · 2000
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y · 2015
Earlier work this paper cites.
Open problem: The landscape of the loss surfaces of multilayer networks
Choromanska, A., LeCun, Y., and Ben Arous, G · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Safran, I., and Shamir, O · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D., and Carmon, Y · 2016
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q., and Hein, M · 2017
Cited alongside, same era.
Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions
Panageas, I., and Piliouras, G · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J., and Bahri, Y · 2017
Cited alongside, same era.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D., and Hoffer, E · 2017
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2019
Later among the works it cites.
Spurious valleys in one-hidden-layer neural network optimization landscapes
Venturi, L., Bandeira, A. S., and Bruna, J · 2019
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L., and Bach, F · 2020
Later among the works it cites.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Fehrman, B., Gess, B., and Jentzen, A · 2020
Later among the works it cites.
Topological properties of the set of functions generated by neural networks of fixed size
Petersen, P., Raslan, M., and Voigtlaender, F · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the power of over-parametrization in neural networks with quadratic activation
Du, S., and Lee, J · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer ReLU neural networks
Safran, I., and Shamir, O · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
Lee, J. D., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M. I., and Recht, B · 2019
Cited alongside, same era.
Optimization and generalization of shallow neural networks with quadratic activation functions
Sarao Mannelli, S., Vanden-Eijnden, E., and Zdeborová, L · 2020
Later among the works it cites.
On the convergence of gradient descent training for two-layer relu-networks in the mean field regime
Wojtowytsch, S · 2020
Later among the works it cites.
Eberle, S., Jentzen, A., Riekert, A., and Weiss, G. S · 2021
Closest in time.
Jentzen, A., and Riekert, A · 2021
Closest in time.
A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
Cheridito, P., Jentzen, A., Riekert, A., and Rossmannek, F · 2022
Closest in time.