Fetching the paper…
Reading the bibliography…
Various gradient compression schemes have been proposed to mitigate the communication cost in distributed training of large scale machine learning models.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The Rprop algorithm
Martin Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann Le Cun · 2003
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
On the absolute constants in the berry–esseen type inequalities for identically distributed summands
Irina Shevtsova · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
ADADELTA: An Adaptive Learning Rate Method
Matthew D. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Stochastic spectral descent for restricted boltzmann machines
David Carlson, Volkan Cevher, and Lawrence Carin · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
GradiVeQ: Vector quantization for bandwidth-efficient gradient aggregation in distributed CNN training
Mingchao Yu, Zhifeng Lin, Krishna Narra, Songze Li, Youjie Li, Nam Sung Kim, Alexander Schwing, Murali Annavaram, and Salman Avestimehr · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
signSGD with majority vote is communication efficient and fault tolerant
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
signSGD via zeroth-order oracle
Sijia Liu, Pin-Yu Chen, Xiangyi Chen, and Mingyi Hong · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin · 2018
Cited alongside, same era.
Atomo: Communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary Charles, Dimitris Papailiopoulos, and Stephen Wright · 2018
Cited alongside, same era.
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Closest in time.
SGD with arbitrary sampling: General analysis and improved rates
Xun Qian, Peter Richtárik, Robert Mansel Gower, Alibek Sailanbayev, Nicolas Loizou, and Egor Shulgin · 2019
Closest in time.
On the convergence of Adam and beyond
Sashank Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Closest in time.
DoubleSqueeze
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Closest in time.
Fast and faster convergence of SGD for over-parameterized models (and an accelerated perceptron)
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Closest in time.
PowerSGD: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Closest in time.
Distributed training with heterogeneous data: Bridging median- and mean-based algorithms
Xiangyi Chen, Tiancong Chen, Haoran Sun, Zhiwei Steven Wu, and Mingyi Hong · 2020
Closest in time.
Momentum improves normalized SGD
Ashok Cutkosky and Harsh Mehta · 2020
Closest in time.