Fetching the paper…
Reading the bibliography…
It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities.
T. Tieleman and G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural Networks for Machine Learning 4 , 2012
2012
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang et al. , “Large scale distributed deep networks,” in Advances in neural information processing systems , 2012, pp. 1223–1231
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
I. Bárány and V. S. Grinberg, “Block partitions of sequences,” Israel Journal of Mathematics , vol. 206, no. 1, pp. 155–164, 2015
2015
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: http://tensorflow.org/
2015
Earlier work this paper cites.
J. Demouth, “Cuda pro tip: Minimize the tail effect,” 2015. [Online]. Available: https://devblogs.nvidia.com/cuda-pro-tip-minimize-the-tail-effect/
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
P. Covington, J. Adams, and E. Sargin, “Deep neural networks for youtube recommendations,” in Proceedings of the 10th ACM Conference on Recommender Systems, ACM, New York, NY, USA . ACM, 2016
2016
Earlier work this paper cites.
A New Lightweight, Modular, and Scalable Deep Learning Framework , 2016, https://caffe2.ai/
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xi, and E. P. Xing, “Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters,” in 2017 USENIX Annual Technical Conference (USENIX ATC 17) . Santa Clara, CA: USENIX Association, Jul. 2017, pp. 181–193. [Online]. Available: https://www.usenix.org/conference/atc17/technical-sessions/presentation/zhang
2017
Earlier work this paper cites.
A. Mirhoseini, H. Pham, Q. V. Le, B. Steiner, R. Larsen, Y. Zhou, N. Kumar, M. Norouzi, S. Bengio, and J. Dean, “Device placement optimization with reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2017, pp. 2430–2439
2017
Earlier work this paper cites.
J. Wang, P. Huang, H. Zhao, Z. Zhang, B. Zhao, and D. L. Lee, “Billion-scale commodity embedding for e-commerce recommendation in alibaba,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . ACM, 2018, pp. 839–848
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles . ACM, 2019, pp. 1–15
2019
Later among the works it cites.
2019
Later among the works it cites.
GpipeTalk , 2019, https://www.youtube.com/watch?v=9s2cum25Kkc
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. A. R. Shah, W. Wu, Q. Lu, L. Zhang, S. Sasidharan, P. DeMar, C. Guok, J. Macauley, E. Pouyoul, J. Kim et al. , “Amoebanet: An sdn-enabled network service for big data science,” Journal of Network and Computer Applications , vol. 119, pp. 70–82, 2018
2018
Cited alongside, same era.
Z. Jia, S. Lin, C. R. Qi, and A. Aiken, “Exploring the hidden dimension in accelerating convolutional neural networks,” 2018
2018
Cited alongside, same era.
Baidu-allreduce , 2018, https://github.com/baidu-research/baidu-allreduce
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. D. Le, T. Sekiyama, Y. Negishi, H. Imai, and K. Kawachiya, “Involving cpus into multi-gpu deep learning,” in Proceedings of the 2018 ACM/SPEC International Conference on Performance Engineering . ACM, 2018, pp. 56–67
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
NVLink , 2019, https://www.nvidia.com/en-us/data-center/nvlink/
2019
Later among the works it cites.
H. Lee, M. Jeong, C. Kim, S. Lim, I. Kim, W. Baek, and B. Yoon, “torchgpipe, A GPipe implementation in PyTorch,” https://github.com/kakaobrain/torchgpipe, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
NCCL , 2019, https://developer.nvidia.com/nccl
2019
Later among the works it cites.
S. Pal, E. Ebrahimi, A. Zulfiqar, Y. Fu, V. Zhang, S. Migacz, D. Nellans, and P. Gupta, “Optimizing multi-gpu parallelization strategies for deep learning training,” IEEE Micro , vol. 39, no. 5, pp. 91–101, 2019
2019
Later among the works it cites.
J. Geng, D. Li, and S. Wang, “Horizontal or vertical?: A hybrid approach to large-scale distributed machine learning,” in Proceedings of the 10th Workshop on Scientific Cloud Computing . ACM, 2019, pp. 1–4
2019
Later among the works it cites.
2019
Later among the works it cites.
Byteps, A high performance and generic framework for distributed DNN training , 2019, https://github.com/bytedance/byteps
2019
Later among the works it cites.
Y. You, J. Hseu, C. Ying, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large-batch training for lstm and beyond,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2019, pp. 1–16
2019
Later among the works it cites.
J. Geng, D. Li, and S. Wang, “Rima: An rdma-accelerated model-parallelized solution to large-scale matrix factorization,” in 2019 IEEE 35th International Conference on Data Engineering (ICDE) . IEEE, 2019, pp. 100–111
2019
Later among the works it cites.
N. Dryden, N. Maruyama, T. Moon, T. Benson, M. Snir, and B. Van Essen, “Channel and filter parallelism for large-scale cnn training,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2019, pp. 1–20
2019
Later among the works it cites.
J. Geng, D. Li, and S. Wang, “Elasticpipe: An efficient and dynamic model-parallel solution to dnn training,” in Proceedings of the 10th Workshop on Scientific Cloud Computing . ACM, 2019, pp. 5–9
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Wang, C.-c. Huang, and J. Li, “Supporting very large models using automatic dataflow graph partitioning,” in Proceedings of the Fourteenth EuroSys Conference 2019 , 2019, pp. 1–17
2019
Later among the works it cites.