Fetching the paper…
Reading the bibliography…
We analyze Local SGD (aka parallel or federated SGD) and Minibatch SGD in the heterogeneous distributed setting, where each machine has access to stochastic gradient estimates for a different, machine-specific, convex objective; the goal is to optimize w.r.t.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Parallel and distributed computation: numerical methods , volume 23
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Constrained consensus and optimization in multi-agent networks
Angelia Nedic, Asuman Ozdaglar, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Distributed stochastic subgradient projection algorithms for convex optimization
S Sundhar Ram, Angelia Nedić, and Venugopal V Veeravalli · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Fan Zhou and Guojing Cong · 2018
Later among the works it cites.
Advances and open problems in federated learning, 2019
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, Adrià Gascón, Badih Ghazi, Phillip B. Gibbons, Marco Gruteser, Zaid Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konečný, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrède Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Özgür, Rasmus Pagh, Mariana Raykova, Hang Qi, Daniel Ramage, Ramesh Raskar, Dawn Song, Weikang Song, Sebastian U. Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramèr, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X. Yu, Han Yu, and Sen Zhao · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tight complexity bounds for optimizing composite objectives
Blake E Woodworth and Nati Srebro · 2016
Cited alongside, same era.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
Lower bounds for finding stationary points i
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2017
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Cited alongside, same era.
Later among the works it cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh · 2019
Later among the works it cites.
Better communication complexity for local sgd
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2019
Later among the works it cites.
Unified optimal analysis of the (stochastic) gradient method
Sebastian U Stich · 2019
Later among the works it cites.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich · 2020
Closest in time.
Is local sgd better than minibatch sgd?
Blake Woodworth, Kumar Kshitij Patel, Sebastian U Stich, Zhen Dai, Brian Bullins, H Brendan McMahan, Ohad Shamir, and Nathan Srebro · 2020
Closest in time.