Fetching the paper…
Reading the bibliography…
Many communication-efficient variants of SGD use gradient quantization schemes.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Adaptive quantization in differential PCM coding of speech
P. Cummiskey, N. S. Jayant, and J. L. Flanagan · 1973
Earlier work this paper cites.
A method of solving a convex programming problem with convergence O ( 1 / k 2 ) O(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Intermediate Calculus
M. H. Protter and C. B. Morrey · 1985
Earlier work this paper cites.
Elements of Information Theory
T. M. Cover and J. A. Thomas · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
R. Bekkerman, M. Bilenko, and J. Langford · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Ré, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
Deep learning with COTS HPC systems
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and A. Ng · 2013
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Project adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue J. Apacible, and K. Kalyanaraman · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Cited alongside, same era.
Asynchronous stochastic convex optimization
J. C. Duchi, S. Chaturapruek, and C. Ré · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
T. Yang, Q. Lin, and Z. Li · 2016
Later among the works it cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Z. Li, R. Tomioka, and M. Vojnovic · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang · 2017
Later among the works it cites.
signSGD: Compressed optimisation for non-convex problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Y. Petuum · 2015
Cited alongside, same era.
Deep learning with elastic averaging SGD
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
Cited alongside, same era.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, and M. Devin · 2016
Cited alongside, same era.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
LQ-Nets: Learned quantization for highly accurate and compact deep neural networks
D. Zhang, J. Yang, D. Ye, and G. Hua · 2018
Later among the works it cites.
NUQSGD: Improved communication efficiency for data-parallel SGD via nonuniform quantization
A. Ramezani-Kebrya, F. Faghri, and D. M. Roy · 2019
Later among the works it cites.
Natural compression for distributed deep learning
S. Horváth, C.-Y Ho, L. Horváth, A. N. Sahu, M. Canini, and P. Richtárik · 2019
Later among the works it cites.
On Empirical Comparisons of Optimizers for Deep Learning
Dami Choi, Christopher J. Shallue, Zachary Nado, Jaehoon Lee, Chris J. Maddison, and George E. Dahl · 2019
Later among the works it cites.
Distributed mean estimation with optimal error bounds
D. Alistarh, S. Ashkboos, and P. Davies · 2020
Closest in time.
Don’t waste your bits! squeeze activations and gradients for deep neural networks via TINYSCRIPT
F. Fu, Y. Hu, Y. He, J. Jiang, Y. Shao, C. Zhang, and B. Cui · 2020
Closest in time.