Fetching the paper…
Reading the bibliography…
We consider the problem of distributed mean estimation (DME), in which $n$ machines are each given a local $d$-dimensional vector $x_v \in \mathbb{R}^d$, and must cooperate to estimate the mean of their inputs $\mu = \frac 1n\sum_{v = 1}^n x_v$, while minimizing total communication cost.
Gesammelte Abhandlungen
Hermann Minkowski · 1911
Earlier work this paper cites.
A note on coverings and packings
C. A. Rogers · 1950
Earlier work this paper cites.
Communication complexity of convex optimization
John N Tsitsiklis and Zhi-Quan Luo · 1987
Earlier work this paper cites.
Lattice quantization
Jerry D. Gibson and Khalid Sayood · 1988
Earlier work this paper cites.
Hadamard Matrices and Their Applications
K. J. Horadam · 2007
Earlier work this paper cites.
The fast johnson–lindenstrauss transform and approximate nearest neighbors
Nir Ailon and Bernard Chazelle · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
The cifar-10 dataset
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Communication lower bounds for statistical estimation problems via a distributed data processing inequality
Mark Braverman, Ankit Garg, Tengyu Ma, Huy L Nguyen, and David P Woodruff · 2016
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
A note on lattice packings via lattice refinements
Martin Henk · 2016
Cited alongside, same era.
Sieving for closest lattice vectors (with preprocessing)
Thijs Laarhoven · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Later among the works it cites.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Atomo: Communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary Charles, Dimitris Papailiopoulos, and Stephen Wright · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
On the estimation of the mean of a random vector
Emilien Joly, Gábor Lugosi, Roberto Imbuzeiro Oliveira, et al · 2017
Cited alongside, same era.
Fully quantized distributed gradient descent
Frederik Künstner · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X Yu, Sanjiv Kumar, and H Brendan McMahan · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Cited alongside, same era.
Venkata Gandikota, Raj Kumar Maity, and Arya Mazumdar · 2019
Later among the works it cites.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2019
Later among the works it cites.
Error feedback fixes SignSGD and other gradient compression schemes
S. P. Karimireddy, Q. Rebjock, S. Stich, and M. Jaggi · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Later among the works it cites.
Nuqsgd: Improved communication efficiency for data-parallel sgd via nonuniform quantization
Ali Ramezani-Kebrya, Fartash Faghri, and Daniel M Roy · 2019
Later among the works it cites.
Powersgd: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Later among the works it cites.
Moniqua: Modulo quantized communication in decentralized SGD
Yucheng Lu and Christopher De Sa · 2020
Closest in time.
Ratq: A universal fixed-length quantizer for stochastic optimization
Prathamesh Mayekar and Himanshu Tyagi · 2020
Closest in time.