Fetching the paper…
Reading the bibliography…
Distributed deep neural network (DDNN) training constitutes an increasingly important workload that frequently runs in the cloud.
A method for unconstrained convex minimization problem with the rate of convergence o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Neurocomputing: Foundations of research
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1988
Earlier work this paper cites.
Estimating available bandwidth using packet pair probing
Ningning Hu and Peter Steenkiste · 2002
Earlier work this paper cites.
Evaluation and characterization of available bandwidth probing techniques
Ningning Hu and Peter Steenkiste · 2003
Earlier work this paper cites.
Optimization of collective communication operations in mpich
Rajeev Thakur, Rolf Rabenseifner, and William Gropp · 2005
Earlier work this paper cites.
Vl2: A scalable and flexible data center network
Albert Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, Dave Maltz, Parveen Patel, and Sudipta Sengupta · 2009
Earlier work this paper cites.
Portland: a scalable fault-tolerant layer 2 data center network fabric
Radhika Niranjan Mysore, Andreas Pamboris, Nathan Farrington, Nelson Huang, Pardis Miri, Sivasankar Radhakrishnan, Vikram Subramanya, and Amin Vahdat · 2009
Earlier work this paper cites.
An architecture for parallel topic models
Alexander Smola and Shravan Narayanamurthy · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Earlier work this paper cites.
More effective distributed ML via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R Ganger, Phillip B Gibbons, et al · 2014
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing · 2014
Earlier work this paper cites.
High-performance distributed ML at scale through parameter server consistency models
Wei Dai, Abhimanu Kumar, Jinliang Wei, Qirong Ho, Garth A. Gibson, and Eric P. Xing · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G. Andersen, Alexander Smola, and Kai Yu · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Cited alongside, same era.
Large scale distributed multiclass logistic regression
Pengtao Xie, Jin Kyu Kim, and Eric P. Xing · 2014
Cited alongside, same era.
Dimmwitted: A study of main-memory statistical analytics
Ce Zhang and Christopher Ré · 2014
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Cited alongside, same era.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke · 2016
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2016
Later among the works it cites.
https://chainer.org/general/2017/02/08/Performance-of-Distributed-Deep-Learning-Using-ChainerMN.html
Performance of distributed deep learning using ChainerMN · 2017
Later among the works it cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inside the social network’s (datacenter) network
Arjun Roy, Hongyi Zeng, Jasmeet Bagga, George Porter, and Alex C. Snoeren · 2015
Cited alongside, same era.
Jupiter rising: A decade of clos topologies and centralized control in google’s datacenter network
Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armistead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, and Amin Vahdat · 2015
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2015
Cited alongside, same era.
Managed communication and consistency for fast data-parallel iterative analytics
Jinliang Wei, Wei Dai, Aurick Qiao, Qirong Ho, Henggang Cui, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Cited alongside, same era.
Revisiting distributed synchronous sgd
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Cited alongside, same era.
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally · 2017
Later among the works it cites.
IncBricks: Toward in-network computation with an in-network cache
Ming Liu, Liang Luo, Jacob Nelson, Luis Ceze, Arvind Krishnamurthy, and Kishore Atreya · 2017
Later among the works it cites.
Deepconfig: Automating data center network topologies management with machine learning
Christopher Streiffer, Huan Chen, Theophilus Benson, and Asim Kadav · 2017
Later among the works it cites.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P. Xing · 2017
Later among the works it cites.
https://aws.amazon.com/mxnet/
Apache mxnet on aws · 2018
Closest in time.
https://docs.microsoft.com/en-us/azure/virtual-machines/windows/sizes-hpc
Azure Windows VM sizes - HPC · 2018
Closest in time.
https://github.com/baidu-research/baidu-allreduce
baidu-research/baidu-allreduce · 2018
Closest in time.
https://cloud.google.com/tpu/
Cloud tpus - ml accelerators for tensorflow — google cloud · 2018
Closest in time.
https://caffe2.ai/docs/distributed-training.html
Distributed training — caffe2 · 2018
Closest in time.
https://azure.microsoft.com/en-us/services/machine-learning-studio/
Machine learning — microsoft azure · 2018
Closest in time.
https://mxnet.incubator.apache.org/faq/cloud.html?highlight=ec2
Mxnet on the cloud — mxnet documentation · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Closest in time.
Optimal message scheduling for aggregation
Leyuan Wang, Mu Li, Edo Liberty, and Alex J Smola · 2018
Closest in time.