Fetching the paper…
Reading the bibliography…
In this paper, it is shown theoretically that spurious local minima are common for deep fully-connected networks and convolutional neural networks (CNNs) with piecewise linear activation functions and datasets that cannot be fitted by linear models.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Region configurations for realizability of lattice piecewise-linear models
J. M. Tarela and M. V. Martínez · 1999
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton · 2012
Earlier work this paper cites.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
M Bianchini and F Scarselli · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Matrix completion has no spurious local minimum
R. Ge, J. D. Lee, and T. Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
The landscape of empirical risk for non-convex losses
S. Mei, Y. Bai, and A. Montanari · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Earlier work this paper cites.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Porcupine neural networks: (almost) all local optima are global
Soheil Feizi, Hamid Javadi, Jesse Zhang, and David Tse · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2017
Earlier work this paper cites.
Learning one-hidden-layer neural networks with landscape design
R. Ge, J. D. Lee, and T. Ma · 2017
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Theory of deep learning ii: Landscape of the empirical risk in deep learning
Qianli Liao and Tomaso Poggio · 2017
Earlier work this paper cites.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Earlier work this paper cites.
Exponentially vanishing suboptimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Simon S. Du and Jason D. Lee · 2018
Cited alongside, same era.
Learning one-hiddenlayer neural networks under general input distributions
Weihao Gao, Ashok Vardhan Makkuva, Sewoong Oh, and Pramod Viswanath · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Large scale structure of neural networks loss landscapes
Stanislav Fort and Stanislaw Jastrzebski · 2019
Later among the works it cites.
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick · 2019
Later among the works it cites.
Deep relu networks have surprisingly few activation patterns
Boris Hanin and David Rolnick · 2019
Later among the works it cites.
Depth with nonlinearity creates no bad local minima in resnets
Kenji Kawaguchi and Yoshua Bengio · 2019
Later among the works it cites.
Elimination of all bad local minima in deep learning
Kenji Kawaguchi and Leslie Pack Kaelbling · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James H. von Brecht · 2018
Cited alongside, same era.
The multilinear structure of relu networks
Thomas Laurent and James H. von Brecht · 2018
Cited alongside, same era.
On the benefit of width for neural networks: Disappearance of bad basins
Dawei Li, Tian Ding, and Ruoyu Sun · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Adding one neuron can eliminate all bad local minima
Shiyu Liang, Ruoyu Sun, Jason D. Lee, and R. Srikant · 2018
Cited alongside, same era.
Understanding the loss surface of neural networks for binary classification
Shiyu Liang, Ruoyu Sun, Yixuan Li, and R. Srikant · 2018
Cited alongside, same era.
Later among the works it cites.
Piecewise strong convexity of neural networks
Tristan Milne · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Q. Nguyen and M. Hein · 2019
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2019
Later among the works it cites.
Theoretical insights into the optimization landscape of overparameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2019
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
Later among the works it cites.
Depth creates no more spurious local minima
Li Zhang · 2019
Later among the works it cites.
Sgd converges to global minimum in deep learning via star-convex path
Yi Zhou, Junjie Yang, Huishuai Zhang, Yingbin Liang, and Vahid Tarokh · 2019
Later among the works it cites.
Low-loss connection of weight vectors: distribution-based approaches
Ivan Anokhin and Dmitry Yarotsky · 2020
Later among the works it cites.
Truth or backpropaganda? an empirical investigation of deep learning theory
M. Goldblum, J. Geiping, A. Schwarzschild, M. Moeller, and T. Goldstein · 2020
Later among the works it cites.
Piecewise linear activations substantially shape the loss surfaces of neural networks
Fengxiang He, Bohan Wang, and Dacheng Tao · 2020
Later among the works it cites.
No spurious local minima in deep quadratic networks
Abbas Kazemipour, Brett Larsen, and Shaul Druckmann · 2020
Later among the works it cites.
Understanding global loss landscape of one-hidden-layer relu networks, part 1: theory
Bo Liu · 2020
Later among the works it cites.
Bounds on over-parameterization for guaranteed existence of descent paths in shallow relu networks
Arsalan Sharifnassab, Saber Salehkaleybar, and S. Jamaloddin Golestani · 2020
Later among the works it cites.
The global landscape of neural networks: An overview
Ruoyu Sun, Dawei Li, Shiyu Liang, Tian Ding, and R Srikant · 2020
Later among the works it cites.
On the number of linear regions of convolutional neural networks
Huan Xiong, Lei Huang, Mengyang Yu, Li Liu, Fan Zhu, and Ling Shao · 2020
Later among the works it cites.
Stochastic gradient descent optimizes overparameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Deforming the loss surface to affect the behaviour of the optimizer
Liangming Chen, Long Jin, Xiujuan Du, Shuai Li, and Mei Liu · 2021
Closest in time.
Ringing relus: Harmonic distortion analysis of nonlinear feedforward networks
Christian H.X. Ali Mehmeti-Gopel, David Hartmann, and Michael Wand · 2021
Closest in time.