Fetching the paper…
Reading the bibliography…
We provide a theoretical algorithm for checking local optimality and escaping saddles at nondifferentiable points of empirical risks of two-layer ReLU networks.
Copositive matrices and definiteness of quadratic forms subject to homogeneous linear inequality constraints
Duncan Henry Martin and David Harris Jacobson · 1981
Earlier work this paper cites.
Some NP-complete problems in quadratic and nonlinear programming
Katta G Murty and Santosh N Kabadi · 1987
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim Blum and Ronald L Rivest · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Quadratic programming with one negative eigenvalue is NP-hard
Panos M Pardalos and Stephen A Vavasis · 1991
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu · 1995
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Xiao-Hu Yu and Guo-An Chen · 1995
Earlier work this paper cites.
Eigenvalue analysis of equilibrium processes defined by linear complementarity conditions
Alberto Seeger · 1999
Earlier work this paper cites.
Nonsmooth analysis and control theory , volume 178
Francis H Clarke, Yuri S Ledyaev, Ronald J Stern, and Peter R Wolenski · 2008
Earlier work this paper cites.
Convex analysis and nonlinear optimization: theory and examples
Jonathan Borwein and Adrian S Lewis · 2010
Earlier work this paper cites.
A variational approach to copositive matrices
J-B Hiriart-Urruty and Alberto Seeger · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Accelerated methods for non-convex optimization
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Globally optimal gradient descent for a ConvNet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
Optimization landscape and expressivity of deep CNNs
Quynh Nguyen and Matthias Hein · 2018
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Closest in time.
Are ResNets provably better than linear predictors?
Ohad Shamir · 2018
Closest in time.
Learning ReLU networks on linearly separable data: Algorithm, optimality, and generalization
Gang Wang, Georgios B Giannakis, and Jie Chen · 2018
Closest in time.
No spurious local minima in a two hidden unit ReLU network
Chenwei Wu, Jiajun Luo, and Jason D Lee · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A generic approach for escaping saddle points
Sashank J Reddi, Manzil Zaheer, Suvrit Sra, Barnabas Poczos, Francis Bach, Ruslan Salakhutdinov, and Alexander J Smola · 2017
Cited alongside, same era.
Learning ReLUs via gradient descent
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Closest in time.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Closest in time.
Learning one-hidden-layer ReLU networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2018
Closest in time.
Critical points of neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
Closest in time.
SGD converges to global minimum in deep learning via star-convex path
Yi Zhou, Junjie Yang, Huishuai Zhang, Yingbin Liang, and Vahid Tarokh · 2019
Closest in time.