Fetching the paper…
Reading the bibliography…
Distributed stochastic gradient descent (SGD) is essential for scaling the machine learning algorithms to a large number of computing nodes.
“The pagerank citation ranking: Bringing order to the web.,”
Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd, · 1999
Earlier work this paper cites.
“The second eigenvalue of the google matrix,”
Taher Haveliwala and Sepandar Kamvar, · 2003
Earlier work this paper cites.
“Learning multiple layers of features from tiny images,”
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton, · 2009
Earlier work this paper cites.
“On the importance of initialization and momentum in deep learning,”
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Scaling distributed machine learning with the parameter server.,”
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su, · 2014
Earlier work this paper cites.
“Deep learning with elastic averaged SGD,”
S. Zhang, A. Choromanska, and Y. LeCun, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Accurate, large minibatch SGD: Training ImageNet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Earlier work this paper cites.
“Communication-efficient learning of deep networks from decentralized data,”
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas, · 2017
Earlier work this paper cites.
“Pytorch: Tensors and dynamic neural networks in python with strong gpu acceleration,” 2017
Adam Paszke, Soumith Chintala, Ronan Collobert, Koray Kavukcuoglu, Clement Farabet, Samy Bengio, Iain Melvin, Jason Weston, and Johnny Mariethoz, · 2017
Cited alongside, same era.
“Slow and stale gradients can win the race: Error-runtime trade-offs in distributed SGD,”
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar, · 2018
Cited alongside, same era.
“On the convergence properties of a k k -step averaging stochastic gradient descent algorithm for nonconvex optimization,”
Fan Zhou and Guojing Cong, · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi, · 2018
Cited alongside, same era.
“Adaptive communication strategies to achieve the best error-runtime trade-off in local-update SGD,”
“PowerSGD: Practical low-rank gradient compression for distributed optimization,”
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi, · 2019
Later among the works it cites.
“Anytime minibatch: Exploiting stragglers in online distributed optimization,”
Nuwan Ferdinand, Haider Al-Lawati, Stark Draper, and Matthew Nokelby, · 2019
Later among the works it cites.
“Computation scheduling for distributed machine learning with straggling workers,”
Mohammad Mohammadi Amiri and Deniz Gündüz, · 2019
Later among the works it cites.
“Local SGD converges fast and communicates little,”
Sebastian U Stich, · 2019
Later among the works it cites.
“Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning,”
Hao Yu, Sen Yang, and Shenghuo Zhu, · 2019
Later among the works it cites.
“Stochastic gradient push for distributed deep learning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jianyu Wang and Gauri Joshi, · 2018
Cited alongside, same era.
“Network topology and communication-computation tradeoffs in decentralized optimization,”
Angelia Nedić, Alex Olshevsky, and Michael G. Rabbat, · 2018
Cited alongside, same era.
“Scaling neural machine translation,”
Myle Ott, Grangier David Edunov, Sergey, and Michael Auli, · 2018
Cited alongside, same era.
“RoBERTa: A robustly optimized BERT pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Cited alongside, same era.
“Language models are unsupervised multi-task learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Cited alongside, same era.
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat, · 2019
Later among the works it cites.
“SlowMo: Improving communication-efficient distributed SGD with slow momentum,”
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat, · 2019
Later among the works it cites.
“Faster distributed deep net training: Computation and communication decoupled stochastic gradient descent,”
Shuheng Shen, Linli Xu, Jingchang Liu, Xianfeng Liang, and Yifei Cheng, · 2019
Later among the works it cites.
“Osp: Overlapping computation and communication in parameter server for fast machine learning,”
Haozhao Wang, Song Guo, and Ruixuan Li, · 2019
Later among the works it cites.