Fetching the paper…
Reading the bibliography…
Communication bottleneck has been identified as a significant issue in distributed optimization of large-scale learning models.
A stochastic approximation method
Robbins Herbert and Sutton Monro · 1951
Earlier work this paper cites.
On the design of gradient algorithms for digitally implemented adaptive filters
R. Gitlin, J. Mazo, and M. Taylor · 1973
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: the rprop algorithm
M. Riedmiller and H. Braun · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for SVM
Shai Shalev-Shwartz, Yoram Singer, and Nathan Srebro · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis R. Bach and Eric Moulines · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Ré, Stephen J. Wright, and Feng Niu · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2012
Earlier work this paper cites.
RMSprop. Coursera: Neural Networks for Machine Learning, Lecture 6.5
T. Tieleman and G Hinton · 2012
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Y. Zhang, J. C. Duchi, M. I. Jordan, and M. J. Wainwright · 2013
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Y. Zhang, J. C. Duchi, and M. J. Wainwright · 2013
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
Gregory F. Coppola · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Top-k multiclass SVM
Maksim Lapin, Matthias Hein, and Bernt Schiele · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. A. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Kai Chen and Qiang Huo · 2016
Cited alongside, same era.
Deep residual learning for image recognition
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2018
Later among the works it cites.
SGD and hogwild! convergence without the bounded gradients assumption
Lam M. Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takác · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in tensorflow
A. Sergeev and M. D. Balso · 2018
Later among the works it cites.
Sparsified SGD with memory
S. U. Stich, J. B. Cordonnier, and M. Jaggi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Parallel SGD: when does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
QSGD: communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
Stochastic, distributed and federated optimization for machine learning
Jakub Konecný · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
H. Mania, X. Pan, D. S. Papailiopoulos, B. Recht, K. Ramchandran, and M. I. Jordan · 2017
Cited alongside, same era.
Error compensated quantized SGD and its applications to large-scale distributed optimization
J. Wu, W. Huang, J. Huang, and T. Zhang · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
ATOMO: communication-efficient learning via atomic sparsification
H. Wang, S. Sievert, S. Liu, Z. B. Charles, D. S. Papailiopoulos, and S. Wright · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
J. Wangni, J. Wang, J. Liu, and T. Zhang · 2018
Later among the works it cites.
Decentralized consensus optimization with asynchrony and delays
Tianyu Wu, Kun Yuan, Qing Ling, Wotao Yin, and Ali H. Sayed · 2018
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, and Martin Jaggi · 2019
Closest in time.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian U. Stich, and Martin Jaggi · 2019
Closest in time.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Closest in time.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Yang · 2019
Closest in time.
Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Closest in time.