Fetching the paper…
Reading the bibliography…
Local search heuristics for non-convex optimizations are popular in applied machine learning.
A systematic approach to higher-order necessary conditions in optimization theory
Dennis S Bernstein · 1984
Earlier work this paper cites.
Higher order conditions with and without lagrange multipliers
Jack Warga · 1986
Earlier work this paper cites.
Some np-complete problems in quadratic and nonlinear programming
Katta G Murty and Santosh N Kabadi · 1987
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, Manfred K Warmuth, et al · 1996
Earlier work this paper cites.
Learning linear transformations
Alan Frieze, Mark Jerrum, and Ravi Kannan · 1996
Earlier work this paper cites.
Squared functional systems and optimization problems
Yurii Nesterov · 2000
Earlier work this paper cites.
Distributional and lˆ q norm inequalities for polynomials over convex bodies in rˆ n
Anthony Carbery and James Wright · 2001
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Rich Caruana Steve Lawrence Lee Giles · 2001
Earlier work this paper cites.
On-line learning theory of soft committee machines with correlated hidden units–steepest gradient descent and natural gradient descent–
Masato Inoue, Hyeyoung Park, and Masato Okada · 2003
Cited alongside, same era.
Singularities affect dynamics of learning in neuromanifolds
Shun-Ichi Amari, Hyeyoung Park, and Tomoko Ozeki · 2006
Cited alongside, same era.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Cited alongside, same era.
Dynamics of learning near singularities in layered networks
Haikun Wei, Jun Zhang, Florent Cousseau, Tomoko Ozeki, and Shun-ichi Amari · 2008
Cited alongside, same era.
Estimate sequence methods: extensions and approximations
Michel Baes · 2009
Cited alongside, same era.
Structure from local optima: Learning subspace juntas via higher order pca
Most tensor problems are np-hard
Christopher J Hillar and Lek-Heng Lim · 2013
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
On the computational complexity of membership problems for the completely positive cone and its dual
Peter JC Dickinson and Luuk Gijben · 2014
Later among the works it cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Later among the works it cites.
The hierarchy of local minimums in polynomial optimization
Jiawang Nie · 2015
Later among the works it cites.
On the quality of the initial basin in overspecified neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Santosh S Vempala and Ying Xiao · 2011
Cited alongside, same era.
Complexity of random smooth functions on the high-dimensional sphere
Antonio Auffinger, Gerard Ben Arous, et al · 2013
Cited alongside, same era.
The number of eigenvalues of a tensor
Dustin Cartwright and Bernd Sturmfels · 2013
Cited alongside, same era.
Itay Safran and Ohad Shamir · 2015
Later among the works it cites.
When are nonconvex problems not scary?
Ju Sun, Qing Qu, and John Wright · 2015
Later among the works it cites.
Gradient descent converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Closest in time.