Fetching the paper…
Reading the bibliography…
We consider speeding up stochastic gradient descent (SGD) by parallelizing it across multiple workers.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan Mcdonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S Mann · 2009
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Yuchen Zhang, John Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C Duchi, and Martin J Wainwright · 2013
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Earlier work this paper cites.
On data dependence in distributed stochastic optimization
Avleen S Bijral, Anand D Sarwate, and Nathan Srebro · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
On the optimality of averaging in distributed statistical learning
Jonathan D Rosenblatt and Boaz Nadler · 2016
Earlier work this paper cites.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Cited alongside, same era.
Local sgd with periodic averaging: Tighter analysis and adaptive synchronization
Farzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, and Viveck Cadambe · 2019
Later among the works it cites.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2019
Later among the works it cites.
Sebastian U Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Later among the works it cites.
Matcha: Speeding up decentralized sgd via matching decomposition sampling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Adaptive communication strategies to achieve the best error-runtime trade-off in local-update sgd
Jianyu Wang and Gauri Joshi · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Cited alongside, same era.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Fan Zhou and Guojing Cong · 2018
Cited alongside, same era.
Communication trade-offs for local-sgd with large step size
Aymeric Dieuleveut and Kumar Kshitij Patel · 2019
Cited alongside, same era.
An accelerated decentralized stochastic proximal algorithm for finite sums
Hadrien Hendrikx, Francis Bach, and Laurent Massoulié · 2019
Cited alongside, same era.
Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang, Gauri Joshi, and Soummya Kar · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Later among the works it cites.
On the rates of convergence of parallelized averaged stochastic gradient algorithms
Antoine Godichon-Baggioni and Sofiane Saadane · 2020
Closest in time.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich · 2020
Closest in time.
Tighter theory for local sgd on identical and heterogeneous data
A Khaled, K Mishchenko, and P Richtárik · 2020
Closest in time.
Robust asynchronous stochastic gradient-push: asymptotically optimal and network-independent performance for strongly convex functions
Artin Spiridonoff, Alex Olshevsky, and Ioannis Ch Paschalidis · 2020
Closest in time.
Communication-efficient distributed deep learning: A comprehensive survey
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li · 2020
Closest in time.
Is local sgd better than minibatch sgd?
Blake Woodworth, Kumar Kshitij Patel, Sebastian U Stich, Zhen Dai, Brian Bullins, H Brendan McMahan, Ohad Shamir, and Nathan Srebro · 2020
Closest in time.