Fetching the paper…
Reading the bibliography…
The present paper develops a novel aggregated gradient approach for distributed machine learning that adaptively compresses the gradient communication.
Computer Networks: A Systems Approach
Larry L Peterson and Bruce S Davie · 2007
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Sensor-centric data reduction for estimation with WSNs via censoring and quantization
Eric J Msechu and Georgios B Giannakis · 2011
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity gpu cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Deep learning with elastic averaging SGD
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Earlier work this paper cites.
DiSCO: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
On the convergence rate of incremental aggregated gradient algorithms
Mert Gurbuzbalaban, Asuman Ozdaglar, and Pablo A Parrilo · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Randomized distributed mean estimation: Accuracy vs communication
Jakub Konečnỳ and Peter Richtárik · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2018
Later among the works it cites.
Network topology and communication-computation tradeoffs in decentralized optimization
Angelia Nedić, Alex Olshevsky, and Michael Rabbat · 2018
Later among the works it cites.
Sparsified SGD with memory
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Atomo: Communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary Charles, Dimitris Papailiopoulos, and Stephen Wright · 2018
Later among the works it cites.
Cooperative SGD: A unified framework for the design and analysis of communication-efficient SGD algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Cited alongside, same era.
SignSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
LAG: Lazily aggregated gradient for communication-efficient distributed learning
Tianyi Chen, Georgios Giannakis, Tao Sun, and Wotao Yin · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Cited alongside, same era.
Communication-efficient distributed statistical inference
Michael I Jordan, Jason D Lee, and Yun Yang · 2018
Cited alongside, same era.
Efficient decentralized deep learning by dynamic model averaging
Michael Kamp, Linara Adilova, Joachim Sicking, Fabian Hüger, Peter Schlicht, Tim Wirtz, and Stefan Wrobel · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
Sindri Magnússon, Hossein Shokri-Ghadikolaei, and Na Li · 2019
Closest in time.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Closest in time.
On the computation and communication complexity of parallel SGD with dynamic batch sizes for stochastic non-convex optimization
Hao Yu and Rong Jin · 2019
Closest in time.