Fetching the paper…
Reading the bibliography…
Large-scale machine learning models are often trained by parallel stochastic gradient descent algorithms.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Urban Stich, and Martin Jaggi · 1901
Earlier work this paper cites.
Universal codeword sets and representations of the integers
P. Elias · 1975
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Efficient distributed learning with sparsity
Communication compression for decentralized training
Hanlin Tang, Shaoduo Gan, Ce Zhang, Tong Zhang, and Ji Liu · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Later among the works it cites.
ImageNet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Later among the works it cites.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jialei Wang, Mladen Kolar, Nathan Srebro, and Tong Zhang · 2017
Cited alongside, same era.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
SIGNSGD: compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
SGD and hogwild! Convergence without the bounded gradients assumption
Lam Nguyen, PHUONG HA NGUYEN, Marten van Dijk, Peter Richtarik, Katya Scheinberg, and Martin Takac · 2018
Cited alongside, same era.
Sparsified SGD with memory
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 2019
Closest in time.
Communication-efficient distributed statistical inference
Michael I Jordan, Jason D Lee, and Yun Yang · 2019
Closest in time.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Closest in time.
DoubleSqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Closest in time.