Fetching the paper…
Reading the bibliography…
In this paper, we propose a unified analysis of variants of distributed SGD with arbitrary compressions and delayed updates.
Decentralized deep learning with arbitrary communication compression
A. Koloskova, T. Lin, S. U. Stich, and M. Jaggi · 1907
Earlier work this paper cites.
Unified optimal analysis of the (stochastic) gradient method
S. U. Stich · 1907
Earlier work this paper cites.
Scaffold: Stochastic controlled averaging for federated learning
S. P. Karimireddy, S. Kale, M. Mohri, S. J. Reddi, S. U. Stich, and A. T. Suresh · 1910
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1985
Earlier work this paper cites.
Parallel and distributed computation: numerical methods , volume 23
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
A unified theory of decentralized SGD with changing topology and local updates
A. Koloskova, N. Loizou, S. Boreiri, M. Jaggi, and S. U. Stich · 2003
Earlier work this paper cites.
Unified analysis of stochastic gradient methods for composite convex and smooth optimization
A. Khaled, O. Sebbouh, N. Loizou, R. M. Gower, and P. Richtárik · 2006
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
LIBSVM: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate Newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Earlier work this paper cites.
An asynchronous mini-batch algorithm for regularized stochastic optimization
H. R. Feyzmahdavian, A. Aytekin, and M. Johansson · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
J. Konečný, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu · 2017
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
H. Mania, X. Pan, D. Papailiopoulos, B. Recht, K. Ramchandran, and M. I. Jordan · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Cited alongside, same era.
On the convergence of local descent methods in federated learning
F. Haddadpour and M. Mahdavi · 2019
Later among the works it cites.
Natural compression for distributed deep learning
S. Horváth, C.-Y. Ho, Ľudovít Horváth, A. N. Sahu, M. Canini, and P. Richtárik · 2019
Later among the works it cites.
Stochastic distributed learning with gradient quantization and variance reduction
S. Horváth, D. Kovalev, K. Mishchenko, S. Stich, and P. Richtárik · 2019
Later among the works it cites.
Advances and open problems in federated learning
P. Kairouz, H. B. McMahan, B. Avent, A. Bellet, M. Bennis, A. N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings, et al · 2019
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A tight convergence analysis for stochastic gradient descent with delayed updates
Y. Arjevani, O. Shamir, and N. Srebro · 2018
Cited alongside, same era.
Stochastic quasi-gradient methods: Variance reduction via Jacobian sketching
R. M. Gower, P. Richtárik, and F. Bach · 2018
Cited alongside, same era.
SEGA: Variance reduction via gradient sketching
F. Hanzely, K. Mishchenko, and P. Richtárik · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
S. Khirirat, H. R. Feyzmahdavian, and M. Johansson · 2018
Cited alongside, same era.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
R. Leblond, F. Pedregosa, and S. Lacoste-Julien · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Y. Nesterov · 2018
Cited alongside, same era.
SGD and Hogwild! convergence without the bounded gradients assumption
L. Nguyen, P. H. Nguyen, M. Dijk, P. Richtárik, K. Scheinberg, and M. Takáč · 2018
Cited alongside, same era.
A. Koloskova, S. Stich, and M. Jaggi · 2019
Later among the works it cites.
A double residual compression algorithm for efficient distributed learning
X. Liu, Y. Li, J. Tang, and M. Yan · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
K. Mishchenko, E. Gorbunov, M. Takáč, and P. Richtárik · 2019
Later among the works it cites.
S. U. Stich and S. P. Karimireddy · 2019
Later among the works it cites.
DoubleSqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
H. Tang, C. Yu, X. Lian, T. Zhang, and J. Liu · 2019
Later among the works it cites.
On biased compression for distributed learning
A. Beznosikov, S. Horváth, P. Richtárik, and M. Safaryan · 2020
Closest in time.
A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent
E. Gorbunov, F. Hanzely, and P. Richtárik · 2020
Closest in time.
Tighter theory for local SGD on identical and heterogeneous data
A. Khaled, K. Mishchenko, and P. Richtárik · 2020
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
D. Kovalev, S. Horváth, and P. Richtárik · 2020
Closest in time.
Artemis: tight convergence guarantees for bidirectional compression in federated learning
C. Philippenko and A. Dieuleveut · 2020
Closest in time.
Is local SGD better than minibatch SGD?
B. Woodworth, K. K. Patel, S. U. Stich, Z. Dai, B. Bullins, H. B. McMahan, O. Shamir, and N. Srebro · 2020
Closest in time.