Fetching the paper…
Reading the bibliography…
It is important to scale out deep neural network (DNN) training for reducing model training time.
H. Wang, S. Potluri, M. Luo, A. K. Singh, S. Sur, and D. K. Panda, “Mvapich2-gpu: optimized gpu to gpu communication for infiniband clusters,” Computer Science-Research and Development , vol. 26, no. 3-4, p. 257, 2011
2011
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in neural information processing systems , 2011, pp. 693–701
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le et al. , “Large scale distributed deep networks,” in Advances in Neural Information Processing Systems , 2012, pp. 1223–1231
2012
Earlier work this paper cites.
E. R. Sparks, A. Talwalkar, V. Smith, J. Kottalam, X. Pan, J. Gonzalez, M. J. Franklin, M. I. Jordan, and T. Kraska, “Mli: An api for distributed machine learning,” in Data Mining (ICDM), 2013 IEEE 13th International Conference on . IEEE, 2013, pp. 1187–1192
2013
Earlier work this paper cites.
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and N. Andrew, “Deep learning with COTS HPC systems,” in Proceedings of the 30th international conference on machine learning , 2013, pp. 1337–1345
2013
Earlier work this paper cites.
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing, “More effective distributed ML via a stale synchronous parallel parameter server,” in Advances in neural information processing systems , 2013, pp. 1223–1231
2013
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision , pp. 1–42, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Dean, “Large scale deep learning,” 2014. [Online]. Available: https://research.google.com/people/jeff/CIKM-keynote-Nov2014.pdf
2014
Earlier work this paper cites.
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman, “Project adam: Building an efficient and scalable deep learning training system,” in 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14) . Broomfield, CO: USENIX Association, 2014, pp. 571–582
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14) . Broomfield, CO: USENIX Association, 2014, pp. 583–598
2014
Earlier work this paper cites.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “On parallelizability of stochastic gradient descent for speech dnns,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 235–239
2014
Earlier work this paper cites.
——, “1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns.” in INTERSPEECH , 2014, pp. 1058–1062
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, A. J. Smola, and K. Yu, “Communication efficient distributed machine learning with the parameter server,” in Advances in Neural Information Processing Systems , 2014, pp. 19–27
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu, “Petuum: a new platform for distributed machine learning on big data,” Big Data, IEEE Transactions on , vol. 1, no. 2, pp. 49–67, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
H. Li, A. Kadav, E. Kruus, and C. Ungureanu, “Malt: distributed data-parallelism for existing ml applications,” in Proceedings of the Tenth European Conference on Computer Systems . ACM, 2015, p. 3
2015
Cited alongside, same era.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in International Conference on Machine Learning , 2015, pp. 1737–1746
2015
Cited alongside, same era.
2017
Later among the works it cites.
A. F. Aji and K. Heafield, “Sparse communication for distributed gradient descent,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2017, pp. 440–445
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, p. 436, 2015
2015
Cited alongside, same era.
J. Wei, W. Dai, A. Qiao, Q. Ho, H. Cui, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing, “Managed communication and consistency for fast data-parallel iterative analytics,” in Proceedings of the Sixth ACM Symposium on Cloud Computing . ACM, 2015, pp. 381–394
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng, “Tensorflow: A system for large-scale machine learning,” in 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . GA: USENIX Association, Nov. 2016
2016
Cited alongside, same era.
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing, “Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server,” in Proceedings of the Eleventh European Conference on Computer Systems . ACM, 2016, p. 4
2016
Cited alongside, same era.
F. N. Iandola, M. W. Moskewicz, K. Ashraf, and K. Keutzer, “Firecaffe: near-linear acceleration of deep neural network training on compute clusters,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2592–2600
2016
Cited alongside, same era.
P. Sun, Y. Wen, T. N. B. Duong, and S. Yan, “Timed dataflow: Reducing communication overhead for distributed machine learning systems,” in 2016 IEEE 22nd International Conference on Parallel and Distributed Systems (ICPADS) . IEEE, 2016, pp. 1110–1117
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Later among the works it cites.
W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Peng, H. Zhao, Q. Zhang, F. Yang, and L. Zhou, “Gandiva: Introspective cluster scheduling for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . Carlsbad, CA: USENIX Association, 2018, pp. 595–610
2018
Later among the works it cites.
“Pytorch,” https://pytorch.org , 2018
2018
Later among the works it cites.
L. Luo, J. Nelson, L. Ceze, A. Phanishayee, and A. Krishnamurthy, “Parameter hub: A rack-scale parameter server for distributed deep neural network training,” in Proceedings of the ACM Symposium on Cloud Computing , ser. SoCC ’18. New York, NY, USA: ACM, 2018, pp. 41–54
2018
Later among the works it cites.
2018
Later among the works it cites.
“Baidu-allreduce,” https://github.com/baidu-research/baidu-allreduce , 2018
2018
Later among the works it cites.
“Nccl,” https://github.com/NVIDIA/nccl , 2018
2018
Later among the works it cites.
S. H. Hashemi, S. A. Jyothi, and R. H. Campbell, “Tictac: Accelerating distributed deep learning with communication scheduling,” 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer, “Imagenet training in minutes,” in Proceedings of the 47th International Conference on Parallel Processing . ACM, 2018, p. 1
2018
Later among the works it cites.
2018
Later among the works it cites.
“Byteps, a high performance and generic framework for distributed dnn training,” https://github.com/bytedance/byteps , 2019
2019
Closest in time.