Fetching the paper…
Reading the bibliography…
We investigate the loss surface of neural networks.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Xiao-Hu Yu and Guo-An Chen · 1995
Earlier work this paper cites.
A primer of real analytic functions
Steven G Krantz and Harold R Parks · 2002
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2006
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
Local minima in training of neural networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a ConvNet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Porcupine neural networks:(almost) all local optima are global
Soheil Feizi, Hamid Javadi, Jesse Zhang, and David Tse · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
Global optimality in neural network training
Benjamin D Haeffele and René Vidal · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
Optimization landscape and expressivity of deep CNNs
Quynh Nguyen and Matthias Hein · 2018
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Closest in time.
Are ResNets provably better than linear predictors?
Ohad Shamir · 2018
Closest in time.
Neural networks with finite intrinsic dimension have no spurious valleys
Luca Venturi, Afonso Bandeira, and Joan Bruna · 2018
Closest in time.
Learning ReLU networks on linearly separable data: Algorithm, optimality, and generalization
Gang Wang, Georgios B Giannakis, and Jie Chen · 2018
Closest in time.
No spurious local minima in a two hidden unit ReLU network
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning ReLUs via gradient descent
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Chenwei Wu, Jiajun Luo, and Jason D Lee · 2018
Closest in time.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Closest in time.
Learning one-hidden-layer ReLU networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2018
Closest in time.
Critical points of neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Closest in time.
SGD converges to global minimum in deep learning via star-convex path
Yi Zhou, Junjie Yang, Huishuai Zhang, Yingbin Liang, and Vahid Tarokh · 2019
Closest in time.