Fetching the paper…
Reading the bibliography…
Deep learning frameworks have been widely deployed on GPU servers for deep learning applications in both academia and industry.
D. E. Rumelhart, G. E. Hinton, R. J. Williams et al. , “Learning representations by back-propagating errors,” Cognitive modeling , vol. 5, no. 3, p. 1, 1988
1988
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola, “Parallelized stochastic gradient descent,” in Advances in neural information processing systems , 2010, pp. 2595–2603
2010
Earlier work this paper cites.
C. Nvidia, “NVIDIA CUDA C programming guide,” Nvidia Corporation , vol. 120, no. 18, p. 8, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
S. Zhang, C. Zhang, Z. You, R. Zheng, and B. Xu, “Asynchronous stochastic gradient descent for DNN training,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on . IEEE, 2013, pp. 6660–6663
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Lee, J. K. Kim, X. Zheng, Q. Ho, G. A. Gibson, and E. P. Xing, “On model parallelization and scheduling strategies for distributed machine learning,” in Advances in neural information processing systems , 2014, pp. 2834–2842
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server.” in OSDI , vol. 1, no. 10.4, 2014, p. 3
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Cited alongside, same era.
2017
Closest in time.
2017
Closest in time.
S. Shams, R. Platania, K. Lee, and S.-J. Park, “Evaluation of deep learning frameworks over different HPC architectures,” in Distributed Computing Systems (ICDCS), 2017 IEEE 37th International Conference on . IEEE, 2017, pp. 1389–1396
2017
Closest in time.
H. Kim, H. Nam, W. Jung, and J. Lee, “Performance analysis of CNN frameworks for GPUs,” in Performance Analysis of Systems and Software (ISPASS), 2017 IEEE International Symposium on . IEEE, 2017, pp. 55–64
2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
S. Shi, Q. Wang, P. Xu, and X. Chu, “Benchmarking state-of-the-art deep learning software tools,” in Proceedings of the 7th International Conference on Cloud Computing and Big Data, IEEE, Macau, China , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing, “Geeps: Scalable deep learning on distributed GPUs with a GPU-specialized parameter server,” in Proceedings of the Eleventh European Conference on Computer Systems . ACM, 2016, p. 4
2016
Cited alongside, same era.
A. A. Awan, K. Hamidouche, A. Venkatesh, and D. K. Panda, “Efficient large message broadcast using nccl and cuda-aware mpi for deep learning,” in Proceedings of the 23rd European MPI Users’ Group Meeting . ACM, 2016, pp. 15–22
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S.-X. Zou, C.-Y. Chen, J.-L. Wu, C.-N. Chou, C.-C. Tsao, K.-C. Tung, T.-W. Lin, C.-L. Sung, and E. Y. Chang, “Distributed training large-scale deep architectures,” in International Conference on Advanced Data Mining and Applications . Springer, 2017, pp. 18–32
2017
Closest in time.
A. A. Awan, K. Hamidouche, J. M. Hashmi, and D. K. Panda, “S-caffe: Co-designing MPI runtimes and Caffe for scalable deep learning on modern GPU clusters,” in Proceedings of the 22nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming . ACM, 2017, pp. 193–205
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing, “Poseidon: an efficient communication architecture for distributed deep learning on gpu clusters,” in Proceedings of the 2017 USENIX Conference on Usenix Annual Technical Conference . USENIX Association, 2017, pp. 181–193
2017
Closest in time.