Fetching the paper…
Reading the bibliography…
Due to the success of deep learning to solving a variety of challenging machine learning tasks, there is a rising interest in understanding loss functions for training neural networks from a theoretical aspect.
Linear learning: Landscapes and algorithms
P. Baldi · 1989
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
On the problem of local minima in backpropagation
M. Gori and A. Tesi · 1992
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
X. H. Yu and G. A. Chen · 1995
Earlier work this paper cites.
Complex-valued autoencoders
P. Baldi and Z. Lu · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
S. Li, J. Jiao, Y. Han, and T. Weissman · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Porcupine neural networks: (almost) all local optima are global
S. Feizi, H. Javadi, J. Zhang, and D. Tse · 2017
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C. D. Freeman and J. Bruna · 2017
Cited alongside, same era.
Identity matters in deep learning
M. Hardt and T. Ma · 2017
Cited alongside, same era.
Depth creates no bad local minima
H. T. Lu and K. Kawaguchi · 2017
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Closest in time.
How regularization affects the critical points in linear networks
A. Taghvaei, J. W. Kim, and P. Mehta · 2017
Closest in time.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Y. Tian · 2017
Closest in time.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Arous, and Y. LeCun
Cited in the paper.
Open problem: The landscape of the loss surfaces of multilayer networks
A. Choromanska, Y. LeCun, and G. Arous
Cited in the paper.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L. Bartlett, and I. S. Dhillon · 2017
Closest in time.
Characterization of gradient dominance and regularity conditions for neural networks
Y. Zhou and Y. Liang · 2017
Closest in time.