2017

Global optimality conditions for deep neural networks

Yun, Chulhee, Sra, Suvrit, Jadbabaie, Ali

Understand

We study the error landscape of deep linear and nonlinear neural networks with the squared error loss.

  • Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete.
  • For deep linear networks, we present necessary and sufficient conditions for a critical point of the risk function to be a global minimum.
  • Surprisingly, our conditions provide an efficiently checkable test for global optimality, while such tests are typically intractable in nonconvex optimization.

Built on

  • On the input-output stability of time-varying nonlinear feedback systems part one: Conditions derived using concepts of loop gain, conicity, and positivity

    George Zames · 1966

    Earlier work this paper cites.

  • Some NP-complete problems in quadratic and nonlinear programming

    Katta G Murty and Santosh N Kabadi · 1987

    Earlier work this paper cites.

  • Training a 3-node neural network is NP-complete

    Avrim Blum and Ronald L Rivest · 1988

    Earlier work this paper cites.

  • Neural networks and principal component analysis: Learning from examples without local minima

    Pierre Baldi and Kurt Hornik · 1989

    Earlier work this paper cites.

  • On the local minima free condition of backpropagation learning

    Xiao-Hu Yu and Guo-An Chen · 1995

    Earlier work this paper cites.

  • Noninear Systems

    Hassan K Khalil · 1996

    Earlier work this paper cites.

Similar

  • Imagenet classification with deep convolutional neural networks

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012

    Cited alongside, same era.

  • The loss surfaces of multilayer networks

    Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015

    Cited alongside, same era.

  • Deep learning without poor local minima

    Kenji Kawaguchi · 2016

    Cited alongside, same era.

  • No bad local minima: Data independent training error guarantees for multilayer neural networks

    Original

    Daniel Soudry and Yair Carmon · 2016

    Cited alongside, same era.

  • Diverse neural network learns true target functions

    Original

    Bo Xie, Yingyu Liang, and Le Song · 2016

    Cited alongside, same era.

  • Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun

    Cited in the paper.

  • Identity mappings in deep residual networks

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun

    Cited in the paper.

Then

  • Deep residual networks: Representation and optimization properties, 2017

    Peter Bartlett, Steve Evans, and Phil Long · 2017

    Closest in time.

  • Global optimality in neural network training

    Benjamin D Haeffele and René Vidal · 2017

    Closest in time.

  • Identity matters in deep learning

    Moritz Hardt and Tengyu Ma · 2017

    Closest in time.

  • Depth creates no bad local minima

    Original

    Haihao Lu and Kenji Kawaguchi · 2017

    Closest in time.

  • The loss surface of deep and wide neural networks

    Quynh Nguyen and Matthias Hein · 2017

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…