Fetching the paper…
Reading the bibliography…
Training deep neural networks (DNNs) efficiently is a challenge due to the associated highly nonconvex optimization.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Fonctions convexes duales et points proximaux dans un espace hilbertien
Jean-Jacques Moreau · 1962
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels. In: Les Équations aux dérivées partielles
Stanisław Łojasiewicz · 1963
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Sur la geometrie semi-et sous-analytique
Stanisław Łojasiewicz · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Real algebraic geometry , volume 3
Jacek Bochnak, Michel Coste, and Marie-Francoise Roy · 1998
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Krzysztof Kurdyka · 1998
Earlier work this paper cites.
Variational Analysis
R. Tyrrell Rockafellar and Roger J.-B. Wets · 1998
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen J. Wright · 1999
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Hui Zou and Trevor Hastie · 2005
Cited alongside, same era.
CIFAR-10 (Canadian Institute for Advanced Research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Cited alongside, same era.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Cited alongside, same era.
Rectified linear units improve restricted Boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Cited alongside, same era.
Proximal splitting methods in signal processing
Patrick L. Combettes and Jean-Christophe Pesquet · 2011
Cited alongside, same era.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward-backward splitting, and regularized Gauss-Seidel methods
Hedy Attouch, Jérôme Bolte, and Benar Fux Svaiter · 2013
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable ADMM approach
Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein · 2016
Later among the works it cites.
Efficient training of very deep neural networks for supervised hashing
Ziming Zhang, Yuting Chen, and Venkatesh Saligrama · 2016
Later among the works it cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
Heinz H. Bauschke and Patrick L. Combettes · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
A block coordinate descent method for regularized multiconvex optimization with applications to nonnegative tensor factorization and completion
Yangyang Xu and Wotao Yin · 2013
Cited alongside, same era.
Proximal alternating linearized minimization for nonconvex and nonsmooth problems
Jérôme Bolte, Shoham Sabach, and Marc Teboulle · 2014
Cited alongside, same era.
Distributed optimization of deeply nested systems
Miguel Carreira-Perpiñán and Weiran Wang · 2014
Cited alongside, same era.
The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis
Cited in the paper.
Clark subgradients of stratifiable functions
Jérôme Bolte, Aris Daniilidis, Adrian Lewis, and Masahiro Shiota
Cited in the paper.
Non-convex optimization for machine learning
Prateek Jain and Purushottam Kar · 2017
Later among the works it cites.
Accelerated block coordinate proximal gradients with applications in high dimensional statistics
Tsz Kit Lau and Yuan Yao · 2017
Later among the works it cites.
Convergent block coordinate descent for training tikhonov regularized deep neural networks
Ziming Zhang and Matthew Brand · 2017
Later among the works it cites.
Proximal backpropagation
Thomas Frerix, Thomas Möllenhoff, Michael Moeller, and Daniel Cremers · 2018
Closest in time.