Fetching the paper…
Reading the bibliography…
Communicating information, like gradient vectors, between computing nodes in distributed and federated learning is typically an unavoidable burden, resulting in scalability issues.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U Stich, and Martin Jaggi · 1901
Earlier work this paper cites.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 1904
Earlier work this paper cites.
Natural compression for distributed deep learning
Samuel Horváth, Chen-Yu Ho, Ľudovít Horváth, Atal Narayan Sahu, Marco Canini, and Peter Richtárik · 1905
Earlier work this paper cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 1910
Earlier work this paper cites.
Coding theorems for a discrete source with a fidelity criterion
C.E. Shannon · 1948
Earlier work this paper cites.
A mathematical theory of communication
C.E. Shannon · 1959
Earlier work this paper cites.
Constructive approximation of a ball by polytopes
Martin Kochol · 1994
Earlier work this paper cites.
Covering the sphere by equal spherical balls
Károly Böröczky and Gergely Wintsche · 2003
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Covering spheres with spheres
Ilya Dumer · 2007
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
Ron Bekkerman, Mikhail Bilenko, and John Langford · 2011
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Yuchen Zhang, John Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks
N. Dryden, T. Moon, S. A. Jacobs, and B. V. Essen · 2016
Cited alongside, same era.
Federated learning: strategies for improving communication efficiency
Jakub Konečný, H. Brendan McMahan, Felix Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X. Yu, Sanjiv Kumar, and H. Brendan McMahan · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Rate distortion for model Compression:From theory to practice
Weihao Gao, Yu-Han Liu, Chong Wang, and Sewoong Oh · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Later among the works it cites.
Sebastian U. Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
DoubleSqueeze
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Cited alongside, same era.
Randomized distributed mean estimation: accuracy vs communication
Jakub Konečný and Peter Richtárik · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally · 2018
Cited alongside, same era.
Atomo: Communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary Charles, Dimitris Papailiopoulos, and Stephen Wright · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Cited alongside, same era.
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
PowerSGD: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Later among the works it cites.
Communication-efficient distributed blockwise momentum SGD with error-feedback
Shuai Zheng, Ziyue Huang, and James T. Kwok · 2019
Later among the works it cites.
On biased compression for distributed learning
Alexandre Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Closest in time.
Information-theoretic understanding of population risk improvement with model compression
Yuheng Bu, Weihao Gao, Shaofeng Zou, and Venugopal V. Veeravalli · 2020
Closest in time.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Closest in time.
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtárik · 2020
Closest in time.
Mher Safaryan, Egor Shulgin, and Peter Richtárik · 2020
Closest in time.