Fetching the paper…
Reading the bibliography…
The main theme of this work is a unifying algorithm, \textbf{L}oop\textbf{L}ess \textbf{S}ARAH (L2S) for problems formulated as summation of $n$ individual loss functions.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2004
Earlier work this paper cites.
Probability and random processes for electrical and computer engineers
John A Gubner · 2006
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Semi-stochastic gradient descent methods
Jakub Konecnỳ and Peter Richtárik · 2013
Earlier work this paper cites.
Optimization with first-order surrogate functions
Julien Mairal · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
A lower bound for the optimization of finite sums
Alekh Agarwal and Leon Bottou · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Barzilai-Borwein step size for stochastic gradient descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Cited alongside, same era.
Less than a single pass: Stochastically controlled stochastic gradient
Lihua Lei and Michael Jordan · 2017
Cited alongside, same era.
R-spider: A fast riemannian stochastic optimization algorithm with curvature independent rate
Jingzhao Zhang, Hongyi Zhang, and Suvrit Sra · 2018
Later among the works it cites.
Stochastic nested variance reduction for nonconvex optimization
Dongruo Zhou, Pan Xu, and Quanquan Gu · 2018
Later among the works it cites.
Sharp analysis for nonconvex sgd escaping from saddle points
Cong Fang, Zhouchen Lin, and Tong Zhang · 2019
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Dmitry Kovalev, Samuel Horvath, and Peter Richtarik · 2019
Closest in time.
Estimate sequences for variance-reduced stochastic composite optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Cited alongside, same era.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Léon Bottou · 2018
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Bin Hu, Stephen Wright, and Laurent Lessard · 2018
Cited alongside, same era.
SpiderBoost: A class of faster variance-reduced algorithms for nonconvex optimization
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2018
Cited alongside, same era.
Inexact sarah algorithm for stochastic optimization
Lam M Nguyen, Katya Scheinberg, and Martin Takáč
Cited in the paper.
Andrei Kulunchakov and Julien Mairal · 2019
Closest in time.
On the adaptivity of stochastic gradient-based optimization
Lihua Lei and Michael I Jordan · 2019
Closest in time.
Almost tune-free variance reduction
Bingcong Li, Lingda Wang, and Georgios B Giannakis · 2019
Closest in time.
A class of stochastic variance reduced methods with an adaptive stepsize
Yan Liu, Congying Han, and Tiande Guo · 2019
Closest in time.
Optimal finite-sum smooth non-convex optimization with SARAH
Lam M Nguyen, Marten van Dijk, Dzung T Phan, Phuong Ha Nguyen, Tsui-Wei Weng, and Jayant R Kalagnanam · 2019
Closest in time.
ProxSARAH: An efficient algorithmic framework for stochastic composite nonconvex optimization
Nhan H Pham, Lam M Nguyen, Dzung T Phan, and Quoc Tran-Dinh · 2019
Closest in time.
L-SVRG and L-Katyusha with arbitrary sampling
Xun Qian, Zheng Qu, and Peter Richtarik · 2019
Closest in time.