Fetching the paper…
Reading the bibliography…
The parameter server architecture is prevalently used for distributed deep learning.
A bridging model for parallel computation
L. G. Valiant · 1990
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Adaptive stochastic gradient descent optimisation for image registration
S. Klein, J. P. Pluim, M. Staring, and M. A. Viergever · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
An architecture for parallel topic models
A. Smola and S. Narayanamurthy · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. A. Zinkevich, M. Weimer, A. Smola, and L. Li · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
A. Coates, A. Ng, and H. Lee · 2011
Earlier work this paper cites.
Scalable inference in latent variable models
A. Ahmed, M. Aly, J. Gonzalez, S. Narayanamurthy, and A. J. Smola · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Deep learning with COTS HPC systems
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and N. Andrew · 2013
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Y. Kim · 2014
Earlier work this paper cites.
The CIFAR-10 Dataset
A. Krizhevsky, V. Nair, and G. Hinton · 2014
Earlier work this paper cites.
Scaling Distributed Machine Learning with the Parameter Server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Cited alongside, same era.
A recursive recurrent neural network for statistical machine translation
S. Liu, N. Yang, M. Li, and M. Zhou · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Cited alongside, same era.
Convolutional neural networks for distant speech recognition
P. Swietojanski, A. Ghoshal, and S. Renals · 2014
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Cited alongside, same era.
Using the output embedding to improve language models
O. Press and L. Wolf · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Later among the works it cites.
Lighter-communication Distributed Machine Learning via Sufficient Factor Broadcasting
P. Xie, J. K. Kim, Y. Zhou, Q. Ho, A. Kumar, Y. Yu, and E. Xing · 2016
Later among the works it cites.
https://github.com/baidu-research/tensorflow-allreduce/ , 2017
Baidu-Research/Tensorflow-Allreduce · 2017
Later among the works it cites.
Sparse Communication for Distributed Gradient Descent
A. F. Aji and K. Heafield · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep Learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Managed communication and consistency for fast data-parallel iterative analytics
J. Wei, W. Dai, A. Qiao, Q. Ho, H. Cui, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2015
Cited alongside, same era.
Deep image: Scaling up image recognition
R. Wu, S. Yan, Y. Shan, Q. Dang, and G. Sun · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu · 2015
Cited alongside, same era.
Microsoft computational network toolkit (cntk). 2015
D. Yu and X. Huang · 2015
Cited alongside, same era.
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
Gaia: Geo-Distributed Machine Learning Approaching LAN Speeds
K. Hsieh, A. Harlap, N. Vijaykumar, D. Konomis, G. R. Ganger, P. B. Gibbons, and O. Mutlu · 2017
Later among the works it cites.
Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu · 2017
Later among the works it cites.
Device Placement Optimization with Reinforcement Learning
A. Mirhoseini, H. Pham, Q. V. Le, B. Steiner, R. Larsen, Y. Zhou, N. Kumar, M. Norouzi, S. Bengio, and J. Dean · 2017
Later among the works it cites.
PyTorch: Tensors and dynamic neural networks in Python with strong GPU acceleration, 2017
A. Paszke, S. Gross, S. Chintala, and G. Chanan · 2017
Later among the works it cites.
Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs
S. Shi and X. Chu · 2017
Later among the works it cites.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.
Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing · 2017
Later among the works it cites.
http://www.mellanox.com/blog/2018/03/ethernet-interconnect-consi-derations-machine/-learning-infrastructure/
Top-3 Ethernet Interconnect Considerations for your Machine Learning Infrastructure · 2018
Closest in time.
https://www.nvidia.com/en-us/data-center/tesla-v100/ , 2018
Nvidia Tesla V100 · 2018
Closest in time.
Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in tensorflow
A. Sergeev and M. Del Balso · 2018
Closest in time.