Fetching the paper…
Reading the bibliography…
In this work we analyze the role nonlinear activation functions play at stationary points of dense neural network training problems.
Some NP-complete problems in quadratic and nonlinear programming
Murty, K. G. and Kabadi, S. N · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Conjugate Gradient Type Methods for Ill–Posed Problems
Hanke, M · 1995
Earlier work this paper cites.
Regularization of Inverse Problems , volume 375 of Mathematics and Its Applications
Engl, H. W., Hanke, M., and Neubauer, A · 1996
Earlier work this paper cites.
Nonlinear Programming
Bertsekas, D. P · 1997
Cited alongside, same era.
Rank Deficient and Discrete Ill-Posed Problems: Numerical Aspects of Linear Inversion
Hansen, P. C · 1998
Cited alongside, same era.
Cubic regularization of newton method and its global performance
Nesterov, Y. and Polyak, B. T · 2006
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Efficient approaches for escaping higher order saddle points in non-convex optimization
Anandkumar, A. and Ge, R · 2016
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2017
Later among the works it cites.
Theory i: Deep networks and the curse of dimensionality
Poggio, T. and Liao, Q · 2018
Later among the works it cites.
O’Leary-Roseberry, T., Alger, N., and Ghattas, O · 2019
Later among the works it cites.
Deep learning in high dimension: Neural network expression rates for generalized polynomial chaos expansions in uq
Schwab, C. and Zech, J · 2019
Later among the works it cites.
The global optimization geometry of shallow linear neural networks
Zhu, Z., Soudry, D., Eldar, Y. C., and Wakin, M. B · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I
Cited in the paper.
Accelerated gradient descent escapes saddle points faster than gradient descent
Jin, C., Netrapalli, P., and Jordan, M. I
Cited in the paper.