Fetching the paper…
Reading the bibliography…
Existing Deep Learning frameworks exclusively use either Parameter Server(PS) approach or MPI parallelism.
GPU Computing
John D. Owens, Mike Houston, David Luebke, Simon Green, John E. Stone, and James C. Phillips. 2008 · 2008
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Pitch Patarasuk and Xin Yuan. 2009 · 2008
Earlier work this paper cites.
Optimal Bucket Algorithms for Large MPI Collectives on Torus Interconnects. In Proceedings of the 24th ACM International Conference on Supercomputing
Nikhil Jain and Yogish Sabharwal. 2010 · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011 · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Large Scale Distributed Deep Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng. 2012 · 2012
Earlier work this paper cites.
An evaluation of User-Level Failure Mitigation support in MPI
Wesley Bland, Aurelien Bouteiller, Thomas Hérault, Joshua Hursey, George Bosilca, and Jack J. Dongarra. 2013 · 2013
Earlier work this paper cites.
Deep Learning with COTS HPC Systems. In Proceedings of the 30th International Conference on International Conference on Machine Learning - Volume 28
Adam Coates, Brody Huval, Tao Wang, David J. Wu, Andrew Y. Ng, and Bryan Catanzaro. 2013 · 2013
Earlier work this paper cites.
Project Adam: Building an Efficient and Scalable Deep Learning Training System. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14)
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014 · 2014
Earlier work this paper cites.
Evaluating User-Level Fault Tolerance for MPI Applications. In Proceedings of the 21st European MPI Users’ Group Meeting
Ignacio Laguna, David F. Richards, Todd Gamblin, Martin Schulz, and Bronis R. de Supinski. 2014 · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server. In Advances in Neural Information Processing Systems
Mu Li, David G Andersen, Alex J Smola, and Kai Yu. 2014 · 2014
Earlier work this paper cites.
Lessons Learned Implementing User-Level Failure Mitigation in MPICH. In 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGrid 2015, Shenzhen, China, May 4-7, 2015
Wesley Bland, Huiwei Lu, Sangmin Seo, and Pavan Balaji. 2015 · 2015
Earlier work this paper cites.
Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Sydney, NSW, Australia, August 10-13, 2015
Longbing Cao, Chengqi Zhang, Thorsten Joachims, Geoffrey I. Webb, Dragos D. Margineantu, and Graham Williams (Eds.). 2015 · 2015
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015 · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Cited alongside, same era.
FireCaffe: near-linear acceleration of deep neural network training on compute clusters
Forrest N. Iandola, Khalid Ashraf, Matthew W. Moskewicz, and Kurt Keutzer. 2015 · 2015
Cited alongside, same era.
DeepSpark: Spark-Based Deep Learning Supporting Asynchronous Updates and Caffe Compatibility
Hanjoo Kim, Jaehong Park, Jaehee Jang, and Sungroh Yoon. 2016 · 2016
Later among the works it cites.
Communication-Efficient Learning of Deep Networks from Decentralized Data
H. Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Aguera y Arcas. 2016 · 2016
Later among the works it cites.
Virtualizing Deep Neural Networks for Memory-Efficient Neural Network Design
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler. 2016 · 2016
Later among the works it cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder. 2016 · 2016
Later among the works it cites.
Google Cloud
2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asynchronous parallel stochastic gradient for nonconvex optimization. In Advances in Neural Information Processing Systems
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu. 2015 · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Staleness-aware Async-SGD for Distributed Deep Learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu. 2015b · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, and others. 2016 · 2016
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
L. Bottou, F. E. Curtis, and J. Nocedal. 2016 · 2016
Cited alongside, same era.
GeePS: Scalable Deep Learning on Distributed GPUs with a GPU-specialized Parameter Server. In Proceedings of the Eleventh European Conference on Computer Systems
Henggang Cui, Hao Zhang, Gregory R. Ganger, Phillip B. Gibbons, and Eric P. Xing. 2016 · 2016
Cited alongside, same era.
Distributed Deep Learning Using Synchronous Stochastic Gradient Descent
Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere, Karthikeyan Vaidyanathan, Srinivas Sridharan, Dhiraj D. Kalamkar, Bharat Kaul, and Pradeep Dubey. 2016 · 2016
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
IBM Cloud
2017a · 2017
Later among the works it cites.
IBM Minsky
2017b · 2017
Later among the works it cites.
LSF restarting job
2017b · 2017
Later among the works it cites.
mpi-caffe
2017 · 2017
Later among the works it cites.
MPI Resize
2017 · 2017
Later among the works it cites.
S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters. In Proceedings of the 22Nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming
Ammar Ahmad Awan, Khaled Hamidouche, Jahanzeb Maqbool Hashmi, and Dhabaleswar K. Panda. 2017 · 2017
Later among the works it cites.
Revisiting Distributed Synchronous SGD
Xinghao Pan, Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Józefowicz. 2017 · 2017
Later among the works it cites.