Fetching the paper…
Reading the bibliography…
Recently, local SGD has got much attention and been extensively studied in the distributed learning community to overcome the communication bottleneck problem.
Minibatch vs local sgd for heterogeneous distributed learning
Woodworth, B., Patel, K. K., and Srebro, N · 2006
Earlier work this paper cites.
Mime: Mimicking centralized stochastic algorithms in federated learning
Karimireddy, S. P., Jaggi, M., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Federated optimization: Distributed optimization beyond the datacenter
Konečnỳ, J., McMahan, B., and Ramage, D · 2015
Earlier work this paper cites.
Privacy-preserving deep learning
Shokri, R. and Shmatikov, V · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Allen-Zhu, Z. and Hazan, E · 2016
Earlier work this paper cites.
Natasha 2: Faster non-convex optimization than sgd
Allen-Zhu, Z · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
Lei, L., Ju, C., Chen, J., and Jordan, M. I · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Cited alongside, same era.
Local sgd with periodic averaging: Tighter analysis and adaptive synchronization
Haddadpour, F., Kamani, M. M., Mahdavi, M., and Cadambe, V. R · 2019
Later among the works it cites.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., et al · 2019
Later among the works it cites.
Ssrgd: Simple stochastic recursive gradient descent for escaping saddle points
Li, Z · 2019
Later among the works it cites.
Finite-sum smooth optimization with sarah
Nguyen, L. M., van Dijk, M., Phan, D. T., Nguyen, P. H., Weng, T.-W., and Kalagnanam, J. R · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Lin, T., Stich, S. U., Patel, K. K., and Jaggi, M · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little
Stich, S. U · 2018
Cited alongside, same era.
Stochastic nested variance reduction for nonconvex optimization
Zhou, D., Xu, P., and Gu, Q · 2018
Cited alongside, same era.
On the convergence of local descent methods in federated learning
Haddadpour, F. and Mahdavi, M · 2019
Cited alongside, same era.
Scaffold: Stochastic controlled averaging for federated learning
Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A. T
Cited in the paper.
Stochastic variance reduction for nonconvex optimization
Reddi, S. J., Hefny, A., Sra, S., Poczos, B., and Smola, A
Cited in the paper.
Sharma, P., Kafle, S., Khanduri, P., Bulusu, S., Rajawat, K., and Varshney, P. K · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Yu, H., Yang, S., and Zhu, S · 2019
Later among the works it cites.
Faster non-convex federated learning via global and local momentum
Das, R., Acharya, A., Hashemi, A., Sanghavi, S., Dhillon, I. S., and Topcu, U · 2020
Later among the works it cites.
Tighter theory for local sgd on identical and heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2020
Later among the works it cites.
A unified theory of decentralized sgd with changing topology and local updates
Koloskova, A., Loizou, N., Boreiri, S., Jaggi, M., and Stich, S · 2020
Later among the works it cites.