Fetching the paper…
Reading the bibliography…
High network communication cost for synchronizing gradients and parameters is the well-known bottleneck of distributed training.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
A framework for performance modeling and prediction
Allan Snavely, Laura Carrington, Nicole Wolter, Jesus Labarta, Rosa Badia, and Avi Purkayastha · 2002
Earlier work this paper cites.
Gradient descent with sparsification: an iterative algorithm for sparse recovery with restricted isometry property
Rahul Garg and Rohit Khandekar · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Parallel coordinate descent for l1-regularized loss minimization
Joseph K Bradley, Aapo Kyrola, Danny Bickson, and Carlos Guestrin · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc'aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc V. Le, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Deep learning with cots hpc systems
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Ng Andrew · 2013
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Petuum: A new platform for distributed machine learning on big data
Eric P Xing, Qirong Ho, Wei Dai, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu · 2015
Cited alongside, same era.
Sparknet: Training deep networks in spark
Philipp Moritz, Robert Nishihara, Ion Stoica, and Michael I Jordan · 2015
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Cited alongside, same era.
Song Han, Huizi Mao, and William J Dally · 2015
Cited alongside, same era.
Binarized neural networks
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Later among the works it cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Later among the works it cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Later among the works it cites.
Recurrent neural networks with limited numerical precision
Joachim Ott, Zhouhan Lin, Ying Zhang, Shih-Chii Liu, and Yoshua Bengio · 2016
Later among the works it cites.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X Yu, H Brendan McMahan, and Sanjiv Kumar · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Performance modeling and scalability optimization of distributed deep learning systems
Feng Yan, Olatunji Ruwase, Yuxiong He, and Trishul M. Chilimbi · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Cited alongside, same era.
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Later among the works it cites.
Scaling Distributed Machine Learning with System and Algorithm Co-design
Mu Li · 2017
Closest in time.
Revisiting distributed synchronous sgd
Xinghao Pan, Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2017
Closest in time.
Faster cnns with direct sparse convolutions and guided pruning
J Park, S Li, W Wen, PTP Tang, H Li, Y Chen, and P Dubey · 2017
Closest in time.
Learning intrinsic sparse structures within long short-term memory
Wei Wen, Yuxiong He, Samyam Rajbhandari, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, and Hai Li · 2017
Closest in time.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Closest in time.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Closest in time.