Fetching the paper…
Reading the bibliography…
We provide a new analysis of local SGD, removing unnecessary assumptions and elaborating on the difference between two data regimes: identical and heterogeneous.
First Analysis of Local GD on Heterogeneous Data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 1909
Earlier work this paper cites.
Better Communication Complexity for Local SGD
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 1909
Earlier work this paper cites.
Parallel Gradient Distribution in Unconstrained Optimization
Olvi L. Mangasarian · 1995
Earlier work this paper cites.
Distributed Training Strategies for the Structured Perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Parallel distributed computing using Python
Lisandro D. Dalcin, Rodrigo R. Paz, Pablo A. Kler, and Alejandro Cosimo · 2011
Earlier work this paper cites.
Solving variational Inequalities with Stochastic Mirror-Prox algorithm
Anatoli Juditsky, Arkadi Nemirovski, and Claire Tauvel · 2011
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Yuchen Zhang, John Duchi, Michael I. Jordan, and Martin J. Wainwright · 2013
Earlier work this paper cites.
Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
Gregory F. Coppola · 2015
Earlier work this paper cites.
Federated Learning: Strategies for Improving Communication Efficiency
Jakub Konečný, H. Brendan McMahan, Felix X. Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Communication-Efficient Learning of Deep Networks from Decentralized Data
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Earlier work this paper cites.
Federated Learning for Mobile Keyboard Prediction
Andrew Hard, Kanishka Rao, Rajiv Mathews, Françoise Beaufays, Sean Augenstein, Hubert Eichner, Chloé Kiddon, and Daniel Ramage · 2018
Earlier work this paper cites.
A Linear Speedup Analysis of Distributed Deep Learning with Sparse and Quantized Communication
Peng Jiang and Gagan Agrawal · 2018
Cited alongside, same era.
Federated Optimization in Heterogeneous Networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Cited alongside, same era.
When Edge Meets Learning: Adaptive Control for Resource-Constrained Distributed Machine Learning
Shiqiang Wang, Tiffany Tuor, Theodoros Salonidis, Kin K. Leung, Christian Makaya, Ting He, and Kevin Chan · 2018
Cited alongside, same era.
Variance Reduced Local SGD with Lower Communication Complexity
Xianfeng Liang, Shuheng Shen, Jingchang Liu, Zhen Pan, Enhong Chen, and Yifei Cheng · 2019
Closest in time.
Revisiting Stochastic Extragradient
Konstantin Mishchenko, Dmitry Kovalev, Egor Shulgin, Peter Richtárik, and Yura Malitsky · 2019
Closest in time.
FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization
Amirhossein Reisizadeh, Aryan Mokhtari, Hamed Hassani, Ali Jadbabaie, and Ramtin Pedarsani · 2019
Closest in time.
Pranay Sharma, Prashant Khanduri, Saikiran Bulusu, Ketan Rajawat, and Pramod K. Varshney · 2019
Closest in time.
Local SGD Converges Fast and Communicates Little
Sebastian U. Stich · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fan Zhou and Guojing Cong · 2018
Cited alongside, same era.
Qsparse-local-SGD: Distributed SGD with Quantization, Sparsification and Local Computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Cited alongside, same era.
Reducing Noise in GAN Training with Variance Reduced Extragradient
Tatjana Chavdarova, Gauthier Gidel, François Fleuret, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
Communication trade-offs for Local-SGD with large step size
Aymeric Dieuleveut and Kumar Kshitij Patel · 2019
Cited alongside, same era.
SGD: General Analysis and Improved Rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Cited alongside, same era.
On the Convergence of Local Descent Methods in Federated Learning
Farzin Haddadpour and Mehrdad Mahdavi · 2019
Cited alongside, same era.
Local SGD with Periodic Averaging: Tighter Analysis and Adaptive Synchronization
Farzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, and Viveck Cadambe · 2019
Cited alongside, same era.
SCAFFOLD: Stochastic Controlled Averaging for On-Device Federated Learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 2019
Cited alongside, same era.
Closest in time.
Sebastian U. Stich and Sai Praneeth Karimireddy · 2019
Closest in time.
SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 2019
Closest in time.
Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
Cong Xie, Oluwasanmi Koyejo, Indranil Gupta, and Haibin Lin · 2019
Closest in time.
Distributed Optimization for Over-Parameterized Learning
Chi Zhang and Qianxiao Li · 2019
Closest in time.
On the Convergence of FedAvg on Non-IID Data
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang · 2020
Closest in time.
Don’t Use Large Mini-batches, Use Local SGD
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Closest in time.