Fetching the paper…
Reading the bibliography…
In this paper, we study the performance of a large family of SGD variants in the smooth nonconvex regime.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 1904
Earlier work this paper cites.
Natural compression for distributed deep learning
Samuel Horváth, Chen-Yu Ho, Ludovít Horváth, Atal Narayan Sahu, Marco Canini, and Peter Richtárik · 1905
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Boris T Polyak · 1963
Earlier work this paper cites.
Error bounds and convergence analysis of feasible descent methods: a general approach
Zhi-Quan Luo and Paul Tseng · 1993
Earlier work this paper cites.
Degenerate nonlinear programming with a quadratic growth condition
Mihai Anitescu · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien, and Brian McWilliams · 2015
Earlier work this paper cites.
An optimal randomized incremental gradient method
Guanghui Lan and Yi Zhou · 2015
Earlier work this paper cites.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Federated learning: strategies for improving communication efficiency
Jakub Konečný, H. Brendan McMahan, Felix Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Katyusha: the first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Cited alongside, same era.
A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent
Eduard Gorbunov, Filip Hanzely, and Peter Richtárik · 2019
Later among the works it cites.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Later among the works it cites.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2019
Later among the works it cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
SEGA: variance reduction via gradient sketching
Filip Hanzely, Konstantin Mishchenko, and Peter Richtárik · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Cited alongside, same era.
Random gradient extrapolation for distributed and stochastic optimization
Guanghui Lan and Yi Zhou · 2018
Cited alongside, same era.
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2019
Later among the works it cites.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Dmitry Kovalev, Samuel Horváth, and Peter Richtárik · 2019
Later among the works it cites.
A unified variance-reduced accelerated gradient method for convex optimization
Guanghui Lan, Zhize Li, and Yi Zhou · 2019
Later among the works it cites.
Communication efficient decentralized training with multiple local updates
Xiang Li, Wenhao Yang, Shusen Wang, and Zhihua Zhang · 2019
Later among the works it cites.
SSRGD: Simple stochastic recursive gradient descent for escaping saddle points
Zhize Li · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Later among the works it cites.
ProxSARAH: An efficient algorithmic framework for stochastic composite nonconvex optimization
Nhan H Pham, Lam M Nguyen, Dzung T Phan, and Quoc Tran-Dinh · 2019
Later among the works it cites.
L-SVRG and L-Katyusha with arbitrary sampling
Xun Qian, Zheng Qu, and Peter Richtárik · 2019
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Later among the works it cites.
Federated machine learning: Concept and applications
Qiang Yang, Yang Liu, Tianjian Chen, and Yongxin Tong · 2019
Later among the works it cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2020
Closest in time.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Closest in time.
A fast Anderson-Chebyshev acceleration for nonlinear optimization
Zhize Li and Jian Li · 2020
Closest in time.
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtárik · 2020
Closest in time.