Fetching the paper…
Reading the bibliography…
This paper proposes a stochastic variant of a classic algorithm---the cubic-regularized Newton method [Nesterov and Polyak 2006].
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Natural gradient descent for on-line learning
Magnus Rattray, David Saad, and Shun-ichi Amari · 1998
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
MNIST Handwritten Digit Database, 2010
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive cubic regularisation methods for unconstrained optimization. Part II: worst-case function-and derivative-evaluation complexity
Coralia Cartis, Nicholas Gould, and Philippe Toint · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martin Abadi, Paul Barham, et al · 2016
Cited alongside, same era.
Gradient descent efficiently finds the cubic-regularized non-convex Newton step
Yair Carmon and John Duchi · 2016
A trust region algorithm with a worst-case iteration complexity of 𝒪 ( ϵ − 3 / 2 ) \mathcal{O}(\epsilon^{-3/2}) for nonconvex optimization
Frank E Curtis, Daniel P Robinson, and Mohammadreza Samadi · 2017
Closest in time.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Closest in time.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Closest in time.
Sub-sampled cubic regularization for non-convex optimization
Jonas Moritz Kohler and Aurelien Lucchi · 2017
Closest in time.
A generic approach for escaping saddle points
Sashank J Reddi, Manzil Zaheer, et al · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Accelerated methods for non-convex optimization
Yair Carmon, John Duchi, Oliver Hinder, and Aaron Sidford · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Cited alongside, same era.
Natasha 2: Faster non-convex optimization than SGD
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Complete dictionary recovery over the sphere I: Overview and the geometric picture
Ju Sun, Qing Qu, and John Wright
Cited in the paper.
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright
Cited in the paper.
Jeffrey Regier, Michael I Jordan, and Jon McAuliffe · 2017
Closest in time.
Complexity analysis of second-order line-search algorithms for smooth nonconvex optimization
Clément W Royer and Stephen J Wright · 2017
Closest in time.
Newton-type methods for non-convex optimization under inexact Hessian information
Peng Xu, Farbod Roosta-Khorasani, and Michael W Mahoney · 2017
Closest in time.