Fetching the paper…
Reading the bibliography…
The variance reduction class of algorithms including the representative ones, SVRG and SARAH, have well documented merits for empirical risk minimization problems.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course , volume 87
Yurii Nesterov · 2004
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Semi-stochastic gradient descent methods
Jakub Konecnỳ and Peter Richtárik · 2013
Earlier work this paper cites.
Optimization with first-order surrogate functions
Julien Mairal · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Stochastic proximal gradient descent with acceleration techniques
Atsushi Nitanda · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Cited alongside, same era.
A lower bound for the optimization of finite sums
Alekh Agarwal and Leon Bottou · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Later among the works it cites.
Bin Hu, Stephen Wright, and Laurent Lessard · 2018
Later among the works it cites.
Momentum-based variance reduction in non-convex sgd
Ashok Cutkosky and Francesco Orabona · 2019
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Dmitry Kovalev, Samuel Horvath, and Peter Richtarik · 2019
Closest in time.
Estimate sequences for variance-reduced stochastic composite optimization
Andrei Kulunchakov and Julien Mairal · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Barzilai-Borwein step size for stochastic gradient descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Cited alongside, same era.
Non-convex finite-sum optimization via SCSG methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Cited alongside, same era.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Adaptive step sizes in variance reduction via regularization
Bingcong Li and Georgios B Giannakis · 2019
Closest in time.
A class of stochastic variance reduced methods with an adaptive stepsize
Yan Liu, Congying Han, and Tiande Guo · 2019
Closest in time.
Accelerating mini-batch sarah by step size rules
Zhuang Yang, Zengping Chen, and Cheng Wang · 2019
Closest in time.
On the convergence of SARAH and beyond
Bingcong Li, Meng Ma, and Georgios B Giannakis · 2020
Closest in time.