Fetching the paper…
Reading the bibliography…
Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD.
Cme 302: Numerical linear algebra fall 2005/06 lecture 10
Gene H Golub · 2005
Earlier work this paper cites.
Mime: Mimicking centralized stochastic algorithms in federated learning
Sai Praneeth Karimireddy, Martin Jaggi, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh · 2008
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan Mcdonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon Mann · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex Smola · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data (2016)
H Brendan McMahan, Eider Moore, Daniel Ramage, S Hampson, and B Agüera y Arcas · 2016
Earlier work this paper cites.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Earlier work this paper cites.
Introduction to online convex optimization
Elad Hazan et al · 2016
Earlier work this paper cites.
Federated learning: Collaborative machine learning without centralized training data, Apr 2017
Brendan McMahan and Daniel Ramage · 2017
Earlier work this paper cites.
Collaborative pac learning
Avrim Blum, Nika Haghtalab, Ariel D Procaccia, and Mingda Qiao · 2017
Earlier work this paper cites.
Graph oracle models, lower bounds, and gaps for parallel stochastic optimization
Blake E Woodworth, Jialei Wang, Adam Smith, Brendan McMahan, and Nati Srebro · 2018
Earlier work this paper cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Earlier work this paper cites.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Cited alongside, same era.
Improved algorithms for collaborative pac learning
Huy Nguyen and Lydia Zakynthinou · 2018
Cited alongside, same era.
Advances and open problems in federated learning. corr
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2019
Cited alongside, same era.
Communication trade-offs for local-sgd with large step size
Aymeric Dieuleveut and Kumar Kshitij Patel · 2019
Cited alongside, same era.
Stochastic newton and cubic newton methods with simple local linear-quadratic rates
Dmitry Kovalev, Konstantin Mishchenko, and Peter Richtárik · 2019
Cited alongside, same era.
The min-max complexity of distributed stochastic convex optimization with intermittent communication
Blake E Woodworth, Brian Bullins, Ohad Shamir, and Nathan Srebro · 2021
Later among the works it cites.
Bias-variance reduced local sgd for less heterogeneous federated learning
Tomoya Murata and Taiji Suzuki · 2021
Later among the works it cites.
A stochastic newton algorithm for distributed convex optimization
Brian Bullins, Kumar Kshitij Patel, Ohad Shamir, Nathan Srebro, and Blake E Woodworth · 2021
Later among the works it cites.
Fedchain: Chained algorithms for near-optimal communication cost in federated learning
Charlie Hou, Kiran K Thekumparampil, Giulia Fanti, and Sewoong Oh · 2021
Later among the works it cites.
Fedpage: A fast local stochastic gradient method for communication-efficient federated learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unified optimal analysis of the (stochastic) gradient method
Sebastian U Stich · 2019
Cited alongside, same era.
Tighter theory for local sgd on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Cited alongside, same era.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian Stich · 2020
Cited alongside, same era.
Federated accelerated stochastic gradient descent
Honglin Yuan and Tengyu Ma · 2020
Cited alongside, same era.
Fedsplit: An algorithmic framework for fast federated optimization
Reese Pathak and Martin J Wainwright · 2020
Cited alongside, same era.
On the outsized importance of learning rates in local update methods
Zachary Charles and Jakub Konecny · 2020
Cited alongside, same era.
Lower bounds for finding stationary points i
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2020
Cited alongside, same era.
Haoyu Zhao, Zhize Li, and Peter Richtárik · 2021
Later among the works it cites.
Stem: A stochastic two-sided momentum algorithm achieving near-optimal sample and communication complexities for federated learning
Prashant Khanduri, Pranay Sharma, Haibo Yang, Mingyi Hong, Jia Liu, Ketan Rajawat, and Pramod Varshney · 2021
Later among the works it cites.
On the unreasonable effectiveness of federated averaging with heterogeneous data
Jianyu Wang, Rudrajit Das, Gauri Joshi, Satyen Kale, Zheng Xu, and Tong Zhang · 2022
Later among the works it cites.
Sharp bounds for federated averaging (local sgd) and continuous perspective
Margalit R Glasgow, Honglin Yuan, and Tengyu Ma · 2022
Later among the works it cites.
On-demand sampling: Learning optimally from multiple distributions
Nika Haghtalab, Michael Jordan, and Eric Zhao · 2022
Later among the works it cites.
Towards optimal communication complexity in distributed non-convex optimization
Kumar Kshitij Patel, Lingxiao Wang, Blake Woodworth, Brian Bullins, and Nathan Srebro · 2022
Later among the works it cites.
Federated online and bandit convex optimization
Kumar Kshitij Patel, Lingxiao Wang, Aadirupa Saha, and Nathan Srebro · 2023
Later among the works it cites.
On the effect of defections in federated learning and how to prevent them
Minbiao Han, Kumar Kshitij Patel, Han Shao, and Lingxiao Wang · 2023
Later among the works it cites.
Fedexp: Speeding up federated averaging via extrapolation
Divyansh Jhunjhunwala, Shiqiang Wang, and Gauri Joshi · 2023
Later among the works it cites.