Fetching the paper…
Reading the bibliography…
The stochastic variance-reduced gradient method (SVRG) and its accelerated variant (Katyusha) have attracted enormous attention in the machine learning community in the last few years due to their superior theoretical properties and empirical behaviour on training supervised machine learning models via the empirical risk minimization paradigm.
Problem complexity and method efficiency in optimization
Arkadi Nemirovsky and David B. Yudin · 1983
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A Nemirovski, A Juditsky, G Lan, and A Shapiro · 2009
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Earlier work this paper cites.
Mini-batch primal and dual methods for SVMs
Martin Takáč, Avleen Bijral, Peter Richtárik, and Nathan Srebro · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Finito: A faster, permutable incremental gradient method for Big Data problems
Aaron Defazio, Tiberio Caetano, and Justin Domke · 2014
Earlier work this paper cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2014
Earlier work this paper cites.
Stochastic dual coordinate ascent with adaptive probabilities
Dominik Csiba, Zheng Qu, and Peter Richtárik · 2015
Earlier work this paper cites.
Primal method for ERM with flexible mini-batching schemes and non-convex losses
Dominik Csiba and Peter Richtárik · 2015
Cited alongside, same era.
Variance reduced stochastic gradient descent with neighbors
Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien, and Brian McWilliams · 2015
Cited alongside, same era.
Incremental majorization-minimization optimization with application to large-scale machine learning
Julien Mairal · 2015
Cited alongside, same era.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Zheng Qu, Peter Richtárik, and Tong Zhang · 2015
Cited alongside, same era.
A simple practical accelerated method for finite sums
Aaron Defazio · 2016
Cited alongside, same era.
Stochastic block BFGS: squeezing more curvature out of data
S2GD: Semi-stochastic gradient descent methods
Jakub Konečný and Peter Richtárik · 2017
Later among the works it cites.
Non-convex finite-sum optimization via CSSG methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Later among the works it cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
Improving SAGA via a probabilistic interpolation with gradient descent
Adel Bibi, Alibek Sailanbayev, Bernard Ghanem, Robert Mansel Gower, and Peter Richtárik · 2018
Later among the works it cites.
Randomized block cubic Newton method
Nikita Doikov and Peter Richtárik · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robert Mansel Gower, Donald Goldfarb, and Peter Richtárik · 2016
Cited alongside, same era.
Mini-batch semi-stochastic gradient descent in the proximal setting
Jakub Konečný, Jie Lu, Peter Richtárik, and Martin Takáč · 2016
Cited alongside, same era.
SDNA: Stochastic dual Newton ascent for empirical risk minimization
Zheng Qu, Peter Richtárik, Martin Takáč, and Olivier Fercoq · 2016
Cited alongside, same era.
SDCA without duality, regularization, and individual convexity
Shai Shalev-Shwartz · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Later among the works it cites.
Stochastic quasi-gradient methods: variance reduction via Jacobian sketching
Robert Mansel Gower, Peter Richtárik, and Francis Bach · 2018
Later among the works it cites.
Direct acceleration of SAGA using sampled negative momentum
Kaiwen Zhou · 2018
Later among the works it cites.
A simple stochastic variance reduced algorithm with fast convergence rates
Kaiwen Zhou, Fanhua Shang, and James Cheng · 2018
Later among the works it cites.