Fetching the paper…
Reading the bibliography…
When scaling distributed training, the communication overhead is often the bottleneck.
Mpi: a standard message passing interface
David W Walker and Jack J Dongarra · 1996
Earlier work this paper cites.
Slow learners are fast
Martin Zinkevich, Alexander J. Smola, and John Langford · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew J. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Ré, and Stephen J. Wright · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
Matthew D. Zeiler · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, and Phillipp Koehn · 2013
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna Choromanska, and Yann LeCun · 2014
Earlier work this paper cites.
Scalable distributed dnn training using commodity gpu cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zhang · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Jakub Konevcnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Cited alongside, same era.
Fast asynchronous parallel stochastic gradient descent: A lock-free approach with convergence guarantee
Local sgd converges fast and communicates little
Sebastian U. Stich · 2018
Later among the works it cites.
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Xiang Yang, and Shenghuo Zhu · 2018
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, and Martin Jaggi · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shen-Yi Zhao and Wu-Jun Li · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
signsgd: compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Cited alongside, same era.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U. Stich, and Martin Jaggi · 2018
Cited alongside, same era.
Closest in time.
A generic communication scheduler for distributed dnn training acceleration
Y Peng, Y Zhu, Y Chen, Y Bao, B Yi, C Lan, C Wu, and C Guo · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Benoit Steiner, Zachary DeVito, Soumith Chintala, Sam Gross, Adam Paszke, Francisco Massa, Adam Lerer, Gregory Chanan, Zeming Lin, Edward Yang, Alban Desmaison, Alykhan Tejani, Andreas Kopf, James Bradbury, Luca Antiga, Martin Raison, Natalia Gimelshein, Sasank Chilamkurthy, Trevor Killeen, Lu Fang, and Junjie Bai · 2019
Closest in time.
Adagrad stepsizes: sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Closest in time.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, and Cho-Jui Hsieh · 2019
Closest in time.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Xiang Yang · 2019
Closest in time.
Communication-efficient distributed blockwise momentum sgd with error-feedback
Shuai Zheng, Ziyue Huang, and James T. Kwok · 2019
Closest in time.
A sufficient condition for convergences of adam and rmsprop
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2019
Closest in time.
Cser: Communication-efficient sgd with error reset
Cong Xie, Shuai Zheng, Oluwasanmi O Koyejo, Indranil Gupta, Mu Li, and Haibin Lin · 2020
Closest in time.