Fetching the paper…
Reading the bibliography…
We consider a standard distributed optimisation setting where $N$ machines, each holding a $d$-dimensional function $f_i$, aim to jointly minimise the sum of the functions $\sum_{i = 1}^N f_i (x)$.
Probabilistic computations: Toward a unified measure of complexity
Andrew Chi-Chin Yao · 1977
Earlier work this paper cites.
Communication complexity of convex optimization
John N. Tsitsiklis and Zhi-Quan Luo · 1987
Earlier work this paper cites.
Determinism vs. nondeterminism in multiparty communication complexity
Danny Dolev and Tomás Feder · 1992
Earlier work this paper cites.
Communication Complexity
Eyal Kushilevitz and Noam Nisan · 1996
Earlier work this paper cites.
An optimal lower bound on the communication complexity of Gap-Hamming-Distance
Amit Chakrabarti and Oded Regev · 2011
Earlier work this paper cites.
Lower bounds for number-in-hand multiparty communication complexity, made easy
Jeff M Phillips, Elad Verbin, and Qin Zhang · 2012
Earlier work this paper cites.
A tight bound for set disjointness in the message-passing model
Mark Braverman, Faith Ellen, Rotem Oshman, Toniann Pitassi, and Vinod Vaikuntanathan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Yuchen Zhang, John Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
Project Adam: Building an efficient and scalable deep learning training system
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
On communication cost of distributed statistical estimation and dimensionality
Ankit Garg, Tengyu Ma, and Huy Nguyen · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Fundamental limits of online and distributed algorithms for statistical learning and estimation
Ohad Shamir · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Earlier work this paper cites.
Convex Optimization: Algorithms and Complexity
Sébastien Bubeck · 2015
Cited alongside, same era.
Communication lower bounds for statistical estimation problems via a distributed data processing inequality
Mark Braverman, Ankit Garg, Tengyu Ma, Huy L Nguyen, and David P Woodruff · 2016
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Jakub Konečný, H. Brendan McMahan, Felix X. Yu, Peter Richtarik, Ananda Theertha Suresh, and Dave Bacon · 2016
Cited alongside, same era.
Tight complexity bounds for optimizing composite objectives
Blake E Woodworth and Nati Srebro · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Fully quantized distributed gradient descent
Frederik Künstner · 2017
Qsparse-local-SGD: Distributed SGD with quantization, sparsification, and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Later among the works it cites.
Error feedback fixes signSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian Stich, and Martin Jaggi · 2019
Later among the works it cites.
On maintaining linear convergence of distributed learning and optimization under limited communication
S. Magnússon, H. Shokri-Ghadikolaei, and N. Li · 2019
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian Urban Stich · 2019
Later among the works it cites.
PowerSGD: Practical low-rank gradient compression for distributed optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X Yu, Sanjiv Kumar, and H Brendan McMahan · 2017
Cited alongside, same era.
When distributed computation is communication expensive
David P. Woodruff and Qin Zhang · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cedric Renggli · 2018
Cited alongside, same era.
Gradient compression for communication-limited convex optimization
Sarit Khirirat, Mikael Johansson, and Dan Alistarh · 2018
Cited alongside, same era.
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Later among the works it cites.
Linearly converging error compensated sgd
Eduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, and Peter Richtárik · 2020
Closest in time.
The communication complexity of optimization
Santosh S. Vempala, Ruosong Wang, and David P. Woodruff · 2020
Closest in time.
Overlap local-SGD: An algorithmic approach to hide communication delays in distributed SGD
Jianyu Wang, Hao Liang, and Gauri Joshi · 2020
Closest in time.
Communication-efficient distributed optimization with quantized preconditioners
Foivos Alimisis, Peter Davies, and Dan Alistarh · 2021
Closest in time.
New bounds for distributed mean estimation and variance reduction
Peter Davies, Vijaykrishna Gurunanthan, Niusha Moshrefi, Saleh Ashkboos, and Dan Alistarh · 2021
Closest in time.
Distributed second order methods with fast rates and compressed communication
Rustem Islamov, Xun Qian, and Peter Richtárik · 2021
Closest in time.
Nuqsgd: Provably communication-efficient data-parallel sgd via nonuniform quantization
Ali Ramezani-Kebrya, Fartash Faghri, Ilya Markov, Vitalii Aksenov, Dan Alistarh, and Daniel M Roy · 2021
Closest in time.