Fetching the paper…
Reading the bibliography…
We study distributed optimization methods based on the {\em local training (LT)} paradigm: achieving communication efficiency by performing richer local gradient-based training on the clients before parameter averaging.
Stochastic distributed learning with gradient quantization and variance reduction
S. Horváth, D. Kovalev, K. Mishchenko, S. Stich, and P. Richtárik · 1904
Earlier work this paper cites.
Natural compression for distributed deep learning
S. Horváth, C.-Y. Ho, Ľudovít Horváth, A. N. Sahu, M. Canini, and P. Richtárik · 1905
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course (Applied Optimization)
Y. Nesterov · 2004
Earlier work this paper cites.
LibSVM: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Y. Nesterov · 2013
Earlier work this paper cites.
Mini-batch primal and dual methods for SVMs
M. Takáč, A. Bijral, P. Richtárik, and N. Srebro · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
M. Li, T. Zhang, Y. Chen, and A. J. Smola · 2014
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Earlier work this paper cites.
Parallel training of DNNs with natural gradient and parameter averaging
D. Povey, X. Zhang, and S. Khudanpur · 2015
Earlier work this paper cites.
SparkNet: Training deep networks in Spark
P. Moritz, R. Nishihara, I. Stoica, and M. I. Jordan · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Earlier work this paper cites.
First order methods in optimization
A. Beck · 2017
Earlier work this paper cites.
Practical secure aggregation for privacy-preserving machine learning
K. Bonawitz, V. Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth · 2017
Cited alongside, same era.
Federated learning: Collaborative machine learning without centralized training data
B. McMahan and D. Ramage · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. Agüera y Arcas · 2017
Cited alongside, same era.
Importance sampling for minibatches
D. Csiba and P. Richtárik · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
S. Khirirat, H. R. Feyzmahdavian, and M. Johansson · 2018
Cited alongside, same era.
Semi-cyclic stochastic gradient descent
Y. J. Cho, J. Wang, and G. Joshi · 2020
Later among the works it cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
S. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. Suresh · 2020
Later among the works it cites.
Better theory for SGD in the nonconvex world
A. Khaled and P. Richtárik · 2020
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
A. Khaled, K. Mishchenko, and P. Richtárik · 2020
Later among the works it cites.
Federated learning: challenges, methods, and future directions
T. Li, A. K. Sahu, A. Talwalkar, and V. Smith · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Eichner, T. Koren, H. B. McMahan, N. Srebro, and K. Talwar · 2019
Cited alongside, same era.
On the convergence of local descent methods infederated learning
F. Haddadpour and M. Mahdavi · 2019
Cited alongside, same era.
Nonconvex variance reduced optimization with arbitrary sampling
S. Horváth and P. Richtárik · 2019
Cited alongside, same era.
Advances and open problems in federated learning
P. Kairouz, H. B. McMahan, B. Avent, A. Bellet, M. Bennis, A. N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings, R. G. D’Oliveira, H. Eichner, S. E. Rouayheb, D. Evans, J. Gardner, Z. Garrett, A. Gascón, B. Ghazi, P. B. Gibbons, M. Gruteser, Z. Harchaoui, C. He, L. He, Z. Huo, B. Hutchinson, J. Hsu, M. Jaggi, T. Javidi, G. Joshi, M. Khodak, J. Konečný, A. Korolova, F. Koushanfar, S. Koyejo, T. Lepoint, Y. Liu, P. Mittal, M. Mohri, R. Nock, A. Özgür, R. Pagh, M. Raykova, H. Qi, D. Ramage, R. Raskar, D. Song, W. Song, S. U. Stich, Z. Sun, A. T. Suresh, F. Tramèr, P. Vepakomma, J. Wang, L. Xiong, Z. Xu, Q. Yang, F. X. Yu, H. Yu, and S. Zhao · 2019
Cited alongside, same era.
First analysis of local GD on heterogeneous data
A. Khaled, K. Mishchenko, and P. Richtárik · 2019
Cited alongside, same era.
Communication-efficient local decentralized SGD methods
X. Li, W. Yang, S. Wang, and Z. Zhang · 2019
Cited alongside, same era.
Distributed learning with compressed gradient differences
K. Mishchenko, E. Gorbunov, M. Takáč, and P. Richtárik · 2019
Cited alongside, same era.
From local SGD to local fixed point methods for federated learning
G. Malinovsky, D. Kovalev, E. Gasanov, L. Condat, and P. Richtárik · 2020
Later among the works it cites.
C. Philippenko and A. Dieuleveut · 2020
Later among the works it cites.
Minibatch vs local sgd for heterogeneous distributed learning
B. E. Woodworth, K. K. Patel, and N. Srebro · 2020
Later among the works it cites.
On large-cohort training for federated learning
Z. Charles, Z. Garrett, Z. Huo, S. Shmulyian, and V. Smith · 2021
Later among the works it cites.
Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients
A. Mitra, R. Jaafar, G. Pappas, and H. Hassani · 2021
Later among the works it cites.
Shifted compression framework: Generalizations and improvements
E. Shulgin and P. Richtárik · 2021
Later among the works it cites.
A field guide to federated optimization
J. Wang, Z. Charles, Z. Xu, G. Joshi, H. B. McMahan, B. A. y Arcas, M. Al-Shedivat, G. Andrew, S. Avestimehr, K. Daly, D. Data, S. Diggavi, H. Eichner, A. Gadhikar, Z. Garrett, A. M. Girgis, F. Hanzely, A. Hard, C. He, S. Horvath, Z. Huo, A. Ingerman, M. Jaggi, T. Javidi, P. Kairouz, S. Kale, S. P. Karimireddy, J. Konecny, S. Koyejo, T. Li, L. Liu, M. Mohri, H. Qi, S. J. Reddi, P. Richtárik, K. Singhal, V. Smith, M. Soltanolkotabi, W. Song, A. T. Suresh, S. U. Stich, A. Talwalkar, H. Wang, B. worth, S. Wu, F. X. Yu, H. Yuan, M. Zaheer, M. Zhang, T. Zhang, C. Zheng, C. Zhu, and W. Zhu · 2021
Later among the works it cites.
Sharp bounds for federated averaging (local sgd) and continuous perspective
M. R. Glasgow, H. Yuan, and T. Ma · 2022
Closest in time.
ProxSkip: Yes! Local gradient steps provably lead to communication acceleration! Finally!
K. Mishchenko, G. Malinovsky, S. Stich, and P. Richtárik · 2022
Closest in time.