Understand
We study the error landscape of deep linear and nonlinear neural networks with the squared error loss.
- Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete.
- For deep linear networks, we present necessary and sufficient conditions for a critical point of the risk function to be a global minimum.
- Surprisingly, our conditions provide an efficiently checkable test for global optimality, while such tests are typically intractable in nonconvex optimization.
Built on
On the input-output stability of time-varying nonlinear feedback systems part one: Conditions derived using concepts of loop gain, conicity, and positivity
George Zames · 1966
Earlier work this paper cites.
Some NP-complete problems in quadratic and nonlinear programming
Katta G Murty and Santosh N Kabadi · 1987
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim Blum and Ronald L Rivest · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Xiao-Hu Yu and Guo-An Chen · 1995
Earlier work this paper cites.
Noninear Systems
Hassan K Khalil · 1996
Earlier work this paper cites.
Similar
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.
Then
Deep residual networks: Representation and optimization properties, 2017
Peter Bartlett, Steve Evans, and Phil Long · 2017
Closest in time.
Global optimality in neural network training
Benjamin D Haeffele and René Vidal · 2017
Closest in time.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Closest in time.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Closest in time.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…