Fetching the paper…
Reading the bibliography…
We identify a class of over-parameterized deep neural networks with standard activation functions and cross-entropy loss which provably have no bad local valley, in the sense that from any point in parameter space there exists a continuous path on which the cross-entropy loss is non-increasing and gets arbitrarily close to zero.
Mathematical analysis
T. M. Apostol · 1974
Earlier work this paper cites.
Training a 3-node neural network is np-complete
A. Blum and R. L Rivest · 1989
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. Boser, J.S. Denker, D. Henderson, R.E. Howard, W. Hubbard, and L.D. Jackel · 1990
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
X. Yu and G. Chen · 1995
Earlier work this paper cites.
Exponentially many local minima for single neurons
P. Auer, M. Herbster, and M. K. Warmuth · 1996
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
J. Sima · 2002
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Hena, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2015
Earlier work this paper cites.
The zero set of a real analytic function
B. Mityagin · 2015
Earlier work this paper cites.
Complex powers of analytic functions and meromorphic renormalization in qft
V. D. Nguyen · 2015
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
H. Sedghi and A. Anandkumar · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
A. Gautier, Q. Nguyen, and M. Hein · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified networks
I. Safran and O. Shamir · 2016
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Later among the works it cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Y. Tian · 2017
Later among the works it cites.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2017
Later among the works it cites.
Understanding deep learning requires re-thinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. Bartlett, and I. Dhillon · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Cited alongside, same era.
Global optimality in neural network training
B. D. Haeffele and R. Vidal · 2017
Cited alongside, same era.
Identity matters in deep learning
M. Hardt and T. Ma · 2017
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Maaten, and K. Weinberger · 2017
Cited alongside, same era.
Depth creates no bad local minima
H. Lu and K. Kawaguchi · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Closest in time.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
S. Du, J. Lee, Y. Tian, A. Singh, and B. Póczos · 2018
Closest in time.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2018
Closest in time.
Optimization landscape and expressivity of deep cnns
Q. Nguyen and M. Hein · 2018
Closest in time.
Learning deep models: Critical points and local openness
M. Nouiehed and M. Razaviyayn · 2018
Closest in time.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2018
Closest in time.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, and N. Srebro · 2018
Closest in time.
Spurious valleys in two-layer neural network optimization landscapes
L. Venturi, A. S. Bandeira, and J. Bruna · 2018
Closest in time.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
G. Wang, G. B. Giannakis, and J. Chen · 2018
Closest in time.
Deep neural networks with multi-branch architectures are less non-convex
H. Zhang, J. Shao, and R. Salakhutdinov · 2018
Closest in time.