Fetching the paper…
Reading the bibliography…
We present a unified framework for analyzing local SGD methods in the convex and strongly convex regimes for distributed/federated training of supervised machine learning models.
Error feedback fixes signSGD and other gradient compression schemes
S. P. Karimireddy, Q. Rebjock, S. U. Stich, and M. Jaggi · 1901
Earlier work this paper cites.
Scaffold: Stochastic controlled averaging for on-device federated learning
S. P. Karimireddy, S. Kale, M. Mohri, S. J. Reddi, S. U. Stich, and A. T. Suresh · 1910
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Y. E. Nesterov · 1983
Earlier work this paper cites.
Minibatch vs local SGD for heterogeneous distributed learning
B. Woodworth, K. K. Patel, and N. Srebro · 2006
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Earlier work this paper cites.
DiSCO: Distributed optimization for self-concordant empirical loss
Y. Zhang and X. Lin · 2015
Earlier work this paper cites.
QSGD: Randomized quantization for communication-optimal stochastic gradient descent
D. Alistarh, J. Li, R. Tomioka, and M. Vojnovic · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
J. Konečný, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon · 2016
Earlier work this paper cites.
Federated learning of deep networks using model averaging
H. B. McMahan, E. Moore, D. Ramage, and B. A. y Arcas · 2016
Earlier work this paper cites.
AIDE: Fast and communication efficient distributed optimization
S. J. Reddi, J. Konečný, P. Richtárik, B. Póczós, and A. Smola · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
A. F. Aji and K. Heafield · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Earlier work this paper cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
D. Alistarh, T. Hoefler, M. Johansson, N. Konstantinov, S. Khirirat, and C. Renggli · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Cited alongside, same era.
SEGA: Variance reduction via gradient sketching
F. Hanzely, K. Mishchenko, and P. Richtárik · 2018
Cited alongside, same era.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
D. Kovalev, S. Horváth, and P. Richtárik · 2019
Later among the works it cites.
Variance reduced local SGD with lower communication complexity
X. Liang, S. Shen, J. Liu, Z. Pan, E. Chen, and Y. Cheng · 2019
Later among the works it cites.
A stochastic decoupling method for minimizing the sum of smooth and non-smooth functions
K. Mishchenko and P. Richtárik · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
K. Mishchenko, E. Gorbunov, M. Takáč, and P. Richtárik · 2019
Later among the works it cites.
NUQSGD: Improved communication efficiency for data-parallel SGD via nonuniform quantization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local SGD
T. Lin, S. U. Stich, K. K. Patel, and M. Jaggi · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Y. Nesterov · 2018
Cited alongside, same era.
SGD and Hogwild! convergence without the bounded gradients assumption
L. Nguyen, P. H. Nguyen, M. Dijk, P. Richtárik, K. Scheinberg, and M. Takáč · 2018
Cited alongside, same era.
Local SGD converges fast and communicates little
S. U. Stich · 2018
Cited alongside, same era.
Atomo: Communication-efficient learning via atomic sparsification
H. Wang, S. Sievert, S. Liu, Z. Charles, D. Papailiopoulos, and S. Wright · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
J. Wangni, J. Wang, J. Liu, and T. Zhang · 2018
Cited alongside, same era.
A. Ramezani-Kebrya, F. Faghri, and D. M. Roy · 2019
Later among the works it cites.
Unified optimal analysis of the (stochastic) gradient method
S. U. Stich · 2019
Later among the works it cites.
S. U. Stich and S. P. Karimireddy · 2019
Later among the works it cites.
PowerSGD: Practical low-rank gradient compression for distributed optimization
T. Vogels, S. P. Karimireddy, and M. Jaggi · 2019
Later among the works it cites.
Federated variance-reduced stochastic gradient descent with robustness to byzantine attacks
Z. Wu, Q. Ling, T. Chen, and G. B. Giannakis · 2019
Later among the works it cites.
On biased compression for distributed learning
A. Beznosikov, S. Horváth, P. Richtárik, and M. Safaryan · 2020
Closest in time.
Linearly converging error compensated SGD
E. Gorbunov, D. Kovalev, D. Makarenko, and P. Richtárik · 2020
Closest in time.
Federated learning of a mixture of global and local models
F. Hanzely and P. Richtárik · 2020
Closest in time.
Tighter theory for local SGD on identical and heterogeneous data
A. Khaled, K. Mishchenko, and P. Richtárik · 2020
Closest in time.
A unified theory of decentralized SGD with changing topology and local updates
A. Koloskova, N. Loizou, S. Boreiri, M. Jaggi, and S. U. Stich · 2020
Closest in time.
99% of worker-master communication in distributed optimization is not needed
K. Mishchenko, F. Hanzely, and P. Richtárik · 2020
Closest in time.
FedSplit: An algorithmic framework for fast federated optimization
R. Pathak and M. J. Wainwright · 2020
Closest in time.
Fedpaq: A communication-efficient federated learning method with periodic averaging and quantization
A. Reisizadeh, A. Mokhtari, H. Hassani, A. Jadbabaie, and R. Pedarsani · 2020
Closest in time.
Federated accelerated stochastic gradient descent
H. Yuan and T. Ma · 2020
Closest in time.