Fetching the paper…
Reading the bibliography…
Large-scale distributed training requires significant communication bandwidth for gradient exchange that limits the scalability of multi-node training, and requires expensive high-bandwidth network infrastructure.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Acoustical and environmental robustness in automatic speech recognition
Alejandro Acero · 1990
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Efficient algorithms for all-to-all communications in multiport message-passing systems
Jehoshua Bruck, Ching-Tien Ho, Shlomo Kipnis, Eli Upfal, and Derrick Weathersby · 1997
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Optimization of collective reduction operations
Rolf Rabenseifner · 2004
Earlier work this paper cites.
Introduction to algorithms
Thomas H Cormen · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Project adam: Building an efficient and scalable deep learning training system
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Cited alongside, same era.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Cited alongside, same era.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Cited alongside, same era.
Training and investigating residual nets
S. Gross and M. Wilber · 2016
Later among the works it cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2016
Later among the works it cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Later among the works it cites.
Layer Normalization
J. Lei Ba, J. R. Kiros, and G. E. Hinton · 2016
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Cited alongside, same era.
Sparknet: Training deep networks in spark
Philipp Moritz, Robert Nishihara, Ion Stoica, and Michael I Jordan · 2015
Cited alongside, same era.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Cited alongside, same era.
Scalable distributed dnn training using commodity gpu cloud computing
Nikko Strom · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
Eric P Xing, Qirong Ho, Wei Dai, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu · 2015
Cited alongside, same era.
Qsgd: Randomized quantization for communication-optimal stochastic gradient descent
Dan Alistarh, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Cited alongside, same era.
Image classification with mxnet
Apache · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Later among the works it cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Later among the works it cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Closest in time.
Adacomp: Adaptive residual gradient compression for data-parallel distributed training
Chia-Yu Chen, Jungwook Choi, Daniel Brand, Ankur Agrawal, Wei Zhang, and Kailash Gopalakrishnan · 2017
Closest in time.
Federated learning: Collaborative machine learning without centralized training data, 2017
Google · 2017
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Closest in time.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Closest in time.