Fetching the paper…
Reading the bibliography…
As the size and complexity of models and datasets grow, so does the need for communication-efficient variants of stochastic gradient descent that can be deployed to perform parallel model training.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Principles of pulse code modulation
K. W. Cattermole · 1969
Earlier work this paper cites.
Universal codeword sets and representations of the integers
P. Elias · 1975
Earlier work this paper cites.
A method of solving a convex programming problem with convergence O ( 1 / k 2 ) O(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
R. Bekkerman, M. Bilenko, and J. Langford · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Ré, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
Deep learning with cots hpc systems
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and A. Ng · 2013
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue J. Apacible, and K. Kalyanaraman · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Cited alongside, same era.
Asynchronous stochastic convex optimization
J. C. Duchi, S. Chaturapruek, and C. Ré · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Cited alongside, same era.
Taming the wild: A unified analysis of hogwild-style algorithms
C. M. D. Sa, Ce. Zhang, K. Olukotun, and C. Ré · 2015
Cited alongside, same era.
LogNet: Energy-efficient neural networks using logarithmic computation
E. H. Lee, D. Miyashita, E. Chai, B. Murmann, and S. S. Wong · 2017
Later among the works it cites.
Faster CNNs with direct sparse convolutions and guided pruning
J. Park, S. Li, W. Wen, P. Tang, H. Li, Y. Chen, and P. Dubey · 2017
Later among the works it cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang · 2017
Later among the works it cites.
signSGD: Compressed optimisation for non-convex problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Later among the works it cites.
Loss-aware weight quantization of deep networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scalable distributed DNN training using commodity GPU cloud computing
N. Strom · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Y. Petuum · 2015
Cited alongside, same era.
Deep learning with elastic averaging SGD
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, and M. Devin · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Binarized neural networks
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2016
Cited alongside, same era.
L. Hou and J. T. Kwok · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso · 2018
Later among the works it cites.
Communication compression for decentralized training
H. Tang, S. Gan, C. Zhang, T. Zhang, and J. Liu · 2018
Later among the works it cites.
ATOMO: Communication-efficient learning via atomic sparsification
H. Wang, S. Sievert, Z. Charles, S. Liu, S. Wright, and D. Papailiopoulos · 2018
Later among the works it cites.
A unified analysis of stochastic momentum methods for deep learning
Y Yan, T. Yang, Q. Lin, Z. Li, and Y Yang · 2018
Later among the works it cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2018
Later among the works it cites.
Natural compression for distributed deep learning
S. Horváth, C.-Y Ho, L. Horváth, A. N. Sahu, M. Canini, and P. Richtárik · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
S. P. Karimireddy, Q. Rebjock, S. Stich, and M. Jaggi · 2019
Closest in time.
Dimension-free bounds for low-precision training
Z. Li and C. M. De Sa · 2019
Closest in time.