Fetching the paper…
Reading the bibliography…
Due to the high communication cost in distributed and federated learning problems, methods relying on compression of communicated messages are becoming increasingly popular.
Stochastic distributed learning with gradient quantization and variance reduction
Horváth, S., Kovalev, D., Mishchenko, K., Stich, S., and Richtárik, P · 1904
Earlier work this paper cites.
One method to rule them all: variance reduction for data, parameters and many new methods
Hanzely, F. and Richtárik, P · 1905
Earlier work this paper cites.
Natural compression for distributed deep learning
Horváth, S., Ho, C.-Y., Ľudovít Horváth, Sahu, A. N., Canini, M., and Richtárik, P · 1905
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Nesterov, Y · 2004
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Beck, A. and Teboulle, M · 2009
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Nesterov, Y · 2012
Earlier work this paper cites.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Even faster accelerated coordinate descent using non-uniform sampling
Allen-Zhu, Z., Qu, Z., Richtárik, P., and Yuan, Y · 2016
Earlier work this paper cites.
Federated learning: strategies for improving communication efficiency
Konečný, J., McMahan, H. B., Yu, F., Richtárik, P., Suresh, A. T., and Bacon, D · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
Cited alongside, same era.
Distributed optimization with arbitrary local solvers
Ma, C., Konečný, J., Jaggi, M., Smith, V., Jordan, M. I., Richtárik, P., and Takáč, M · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and Agüera y Arcas, B · 2017
Cited alongside, same era.
SEGA: variance reduction via gradient sketching
Hanzely, F., Mishchenko, K., and Richtárik, P · 2018
First analysis of local GD on heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2019
Later among the works it cites.
A unified variance-reduced accelerated gradient method for convex optimization
Lan, G., Li, Z., and Zhou, Y · 2019
Later among the works it cites.
Federated learning: challenges, methods, and future directions
Li, T., Sahu, A. K., Talwalkar, A., and Smith, V · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
Mishchenko, K., Gorbunov, E., Takáč, M., and Richtárik, P · 2019
Later among the works it cites.
L-SVRG and L-Katyusha with arbitrary sampling
Qian, X., Qu, Z., and Richtárik, P · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributed learning with compressed gradients
Khirirat, S., Feyzmahdavian, H. R., and Johansson, M · 2018
Cited alongside, same era.
Sparsified SGD with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2018
Cited alongside, same era.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Karimireddy, S., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A · 2019
Cited alongside, same era.
Accelerated coordinate descent with arbitrary sampling and best rates for minibatches
Hanzely, F. and Richtárik, P
Cited in the paper.
Local SGD converges fast and communicates little
Stich, S. U · 2019
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2020
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Kovalev, D., Horváth, S., and Richtárik, P · 2020
Closest in time.
A unified analysis of stochastic gradient methods for nonconvex federated optimization
Li, Z. and Richtárik, P · 2020
Closest in time.