Fetching the paper…
Reading the bibliography…
Adam is the important optimization algorithm to guarantee efficiency and accuracy for training many important tasks such as BERT and ImageNet.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2011
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don t care
S. Chaturapruek, J. C. Duchi, and C. Ré · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Asynchronous stochastic gradient descent with delay compensation for distributed deep learning
S. Zheng, Q. Meng, T. Wang, W. Chen, N. Yu, Z. Ma, and T. Liu · 2016
Earlier work this paper cites.
QSGD: Communication-Efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Earlier work this paper cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang · 2017
Cited alongside, same era.
cpSGD: Communication-efficient and differentially-private distributed SGD
N. Agarwal, A. T. Suresh, F. X. X. Yu, S. Kumar, and B. McMahan · 2018
Cited alongside, same era.
signsgd with majority vote is communication efficient and byzantine fault tolerant
J. Bernstein, J. Zhao, K. Azizzadenesheli, and A. Anandkumar · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
P. Jiang and G. Agrawal · 2018
Cited alongside, same era.
Pipe-sgd: A decentralized pipelined sgd framework for distributed deep net training
Y. Li, M. Yu, S. Li, S. Avestimehr, N. S. Kim, and A. Schwing · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Later among the works it cites.
Communication-efficient distributed sgd with sketching
N. Ivkin, D. Rothchild, E. Ullah, V. braverman, I. Stoica, and R. Arora · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, and Y. Liu · 2019
Later among the works it cites.
A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks
S. Shi, Q. Wang, K. Zhao, Z. Tang, Y. Wang, X. Huang, and X. Chu · 2019
Later among the works it cites.
Compressing gradient optimizers via Count-Sketches
R. Spring, A. Kyrillidis, V. Mohan, and A. Shrivastava · 2019
Later among the works it cites.
Communication-efficient distributed learning via lazily aggregated quantized gradients
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards more efficient stochastic decentralized learning: Faster convergence and sparse communication
Z. Shen, A. Mokhtari, T. Zhou, P. Zhao, and H. Qian · 2018
Cited alongside, same era.
Sparsified sgd with memory
S. U. Stich, J.-B. Cordonnier, and M. Jaggi · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2018
Cited alongside, same era.
Gradient sparsification for Communication-Efficient distributed optimization
J. Wangni, J. Wang, J. Liu, and T. Zhang · 2018
Cited alongside, same era.
Communication-Computation efficient gradient coding
M. Ye and E. Abbe · 2018
Cited alongside, same era.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations
D. Basu, D. Data, C. Karakus, and S. Diggavi · 2019
Cited alongside, same era.
Distributed learning over unreliable networks
C. Yu, H. Tang, C. Renggli, S. Kassing, A. Singla, D. Alistarh, C. Zhang, and J. Liu
Cited in the paper.
J. Sun, T. Chen, G. Giannakis, and Z. Yang · 2019
Later among the works it cites.
DoubleSqueeze
H. Tang, C. Yu, X. Lian, T. Zhang, and J. Liu · 2019
Later among the works it cites.
Powersgd: Practical low-rank gradient compression for distributed optimization
T. Vogels, S. P. Karimireddy, and M. Jaggi · 2019
Later among the works it cites.
Communication-efficient distributed blockwise momentum sgd with error-feedback
S. Zheng, Z. Huang, and J. Kwok · 2019
Later among the works it cites.
Decentralized deep learning with arbitrary communication compression
A. Koloskova*, T. Lin*, S. U. Stich, and M. Jaggi · 2020
Closest in time.
Distributed sgd with flexible gradient compression
T. T. Phuong and L. T. Phong · 2020
Closest in time.