Fetching the paper…
Reading the bibliography…
In this paper, we study and analyze the mini-batch version of StochAstic Recursive grAdient algoritHm (SARAH), a method employing the stochastic recursive gradient, for solving empirical loss minimization for the case of nonconvex losses.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Boris T Polyak · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T. Polyak · 1964
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Introductory lectures on convex optimization : a basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Pegasos: Primal estimated sub-gradient solver for SVM
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Optimization with first-order surrogate functions
Julien Mairal · 2013
Cited alongside, same era.
A lower bound for the optimization of finite sums
Alekh Agarwal and Leon Bottou · 2015
Later among the works it cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Later among the works it cites.
Mini-batch semi-stochastic gradient descent in the proximal setting
Jakub Konečný, Jie Liu, Peter Richtárik, and Martin Takáč · 2016
Later among the works it cites.
Stochastic variance reduction for nonconvex optimization
Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczos, and Alexander J. Smola · 2016
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2016
Later among the works it cites.
Natasha: Faster non-convex stochastic optimization via strongly non-convex parameter
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien
Cited in the paper.
A faster, permutable incremental gradient method for big data problems
Aaron Defazio, Justin Domke, and Tibério Caetano
Cited in the paper.
Zeyuan Allen Zhu · 2017
Closest in time.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Closest in time.