Fetching the paper…
Reading the bibliography…
A central challenge to using first-order methods for optimizing nonconvex problems is the presence of saddle points.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Pseudogradient adaptation and training algorithms
BT Poljak and Ya Z Tsypkin · 1973
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Lennart Ljung · 1977
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou · 1991
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A. Pearlmutter · 1994
Earlier work this paper cites.
Introductory Lectures On Convex Optimization: A Basic Course
Yurii Nesterov · 2003
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Stochastic approximation methods for constrained and unconstrained systems , volume 26
Harold Joseph Kushner and Dean S Clark · 2012
Earlier work this paper cites.
Krylov subspace descent for deep learning
Oriol Vinyals and Daniel Povey · 2012
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Minimizing Finite Sums with the Stochastic Average Gradient
Mark W. Schmidt, Nicolas Le Roux, and Francis R. Bach · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Later among the works it cites.
Mini-Batch Semi-Stochastic Gradient Descent in the Proximal Setting
Jakub Konečný, Jie Liu, Peter Richtárik, and Martin Takáč · 2015
Later among the works it cites.
An optimal randomized incremental gradient method
Guanghui Lan and Yi Zhou · 2015
Later among the works it cites.
Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Later among the works it cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alekh Agarwal and Leon Bottou · 2014
Cited alongside, same era.
The loss surface of multilayer networks
Anna Choromanska, Mikael Henaff, Michaël Mathieu, Gérard Ben Arous, and Yann LeCun · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Equilibrated adaptive learning rates for non-convex optimization
Yann Dauphin, Harm de Vries, and Yoshua Bengio · 2015
Cited alongside, same era.
Finding approximate local minima for nonconvex optimization in linear time
Naman Agarwal, Zeyuan Allen Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma
Cited in the paper.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex J Smola · 2015
Later among the works it cites.
Accelerated methods for non-convex optimization
Yair Carmon, John C. Duchi, Oliver Hinder, and Aaron Sidford · 2016
Later among the works it cites.
The power of normalization: Faster evasion of saddle points
Kfir Y. Levy · 2016
Later among the works it cites.
Stochastic variance reduction for nonconvex optimization
Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczos, and Alexander J. Smola · 2016
Later among the works it cites.
Global convergence rate analysis of unconstrained optimization methods based on probabilistic models
C. Cartis and K. Scheinberg · 2017
Closest in time.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Closest in time.