Convergence of the iterates of descent methods for analytic cost functions
Pierre-Antoine Absil, Robert Mahony, and Benjamin Andrews · 2005
Earlier work this paper cites.
Nonsmooth analysis and control theory , volume 178
Francis H Clarke, Yuri S Ledyaev, Ronald J Stern, and Peter R Wolenski · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Curves of descent
Dmitriy Drusvyatskiy, Alexander D Ioffe, and Adrian S Lewis · 2015
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Original
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Low-rank solutions of linear matrix equations via procrustes flow
Original
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Benjamin Recht · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Original
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Identity matters in deep learning
Original
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.