Fetching the paper…
Reading the bibliography…
We study local SGD (also known as parallel SGD and federated averaging), a natural and frequently used stochastic distributed optimization method.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew J. Streeter · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Earlier work this paper cites.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Earlier work this paper cites.
Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
Greg Coppola · 2015
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Cited alongside, same era.
On the optimality of averaging in distributed statistical learning
Jonathan D Rosenblatt and Boaz Nadler · 2016
Cited alongside, same era.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
On the rates of convergence of parallelized averaged stochastic gradient algorithms
Antoine Godichon-Baggioni and Sofiane Saadane · 2017
Graph oracle models, lower bounds, and gaps for parallel stochastic optimization
Blake Woodworth, Jialei Wang, Brendan McMahan, and Nathan Srebro · 2018
Later among the works it cites.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Fan Zhou and Guojing Cong · 2018
Later among the works it cites.
Communication trade-offs for local-sgd with large step size
Aymeric Dieuleveut and Kumar Kshitij Patel · 2019
Later among the works it cites.
Advances and open problems in federated learning, 2019
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, Adrià Gascón, Badih Ghazi, Phillip B. Gibbons, Marco Gruteser, Zaid Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konečný, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrède Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Özgür, Rasmus Pagh, Mariana Raykova, Hang Qi, Daniel Ramage, Ramesh Raskar, Dawn Song, Weikang Song, Sebastian U. Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramèr, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X. Yu, Han Yu, and Sen Zhao · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Praneeth Netrapalli, Sham M Kakade, Rahul Kidambi, and Aaron Sidford · 2017
Cited alongside, same era.
Memory and communication efficient distributed stochastic optimization with minibatch-prox
Jialei Wang, Weiran Wang, and Nathan Srebro · 2017
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Cited alongside, same era.
On the randomized complexity of minimizing a convex quadratic function
Max Simchowitz · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Cited alongside, same era.
Local sgd with periodic averaging: Tighter analysis and adaptive synchronization
Farzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, and Viveck Cadambe
Cited in the paper.
Later among the works it cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh · 2019
Later among the works it cites.
Better communication complexity for local sgd
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2019
Later among the works it cites.
Unified optimal analysis of the (stochastic) gradient method
Sebastian U Stich · 2019
Later among the works it cites.
Sebastian U Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
Lecture notes 1 for optimization methods for large-scale systems, 2019
Lieven Vandenberghe · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Later among the works it cites.