Fetching the paper…
Reading the bibliography…
Distributed Deep Learning (DDL) has rapidly grown its popularity since it helps boost the training performance on high-performance GPU clusters.
S. Sarvotham, R. Riedi, and R. Baraniuk, “Connection-level Analysis and Modeling of Network Traffic,” in Proceedings of the 1st ACM SIGCOMM Workshop on Internet Measurement , ser. IMW ’01, New York, NY, USA, 2001, pp. 99–103
2001
Earlier work this paper cites.
R. Rabenseifner, “Optimization of collective reduction operations,” in Computational Science - ICCS 2004 , M. Bubak, G. D. van Albada, P. M. A. Sloot, and J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2004
2004
Earlier work this paper cites.
R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of collective communication operations in mpich,” The International Journal of High Performance Computing Applications , vol. 19, no. 1, pp. 49–66, 2005
2005
Earlier work this paper cites.
T. Hoefler, W. Gropp, R. Thakur, and J. L. Träff, “Toward performance models of mpi implementations for understanding application scaling issues,” in Recent Advances in the Message Passing Interface , R. Keller, E. Gabriel, M. Resch, and J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010, pp. 21–30
2010
Earlier work this paper cites.
G. L. Stavrinides and H. D. Karatza, “Scheduling multiple task graphs in heterogeneous distributed real-time systems by exploiting schedule holes with bin packing techniques,” Simulation Modelling Practice and Theory , vol. 19, no. 1, pp. 540 – 552, 2011
2011
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. aurelio Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng, “Large Scale Distributed Deep Networks,” in Advances in Neural Information Processing Systems 25 , F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, Eds., 2012, pp. 1223–1231
2012
Earlier work this paper cites.
M. Li, D. G. Andersen, A. J. Smola, and K. Yu, “Communication Efficient Distributed Machine Learning with the Parameter Server,” in Advances in Neural Information Processing Systems 27 , 2014, pp. 19–27
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep Learning,” nature , vol. 521, no. 7553, p. 436, 2015
2015
Earlier work this paper cites.
V. Jalaparti, P. Bodik, I. Menache, S. Rao, K. Makarychev, and M. Caesar, “Network-aware scheduling for data-parallel jobs: Plan when you can,” in Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication , ser. SIGCOMM ’15. New York, NY, USA: ACM, 2015, pp. 407–420
2015
Earlier work this paper cites.
A. Tumanov, T. Zhu, J. W. Park, M. A. Kozuch, M. Harchol-Balter, and G. R. Ganger, “Tetrisched: Global rescheduling with adaptive plan-ahead in dynamic heterogeneous clusters,” in Proceedings of the Eleventh European Conference on Computer Systems , ser. EuroSys ’16. New York, NY, USA: ACM, 2016, pp. 35:1–35:16
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Foley and J. Danskin, “Ultra-performance pascal gpu and nvlink interconnect,” IEEE Micro , vol. 37, no. 2, pp. 7–17, Mar 2017
2017
Earlier work this paper cites.
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing, “Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters,” in 2017 USENIX Annual Technical Conference (USENIX ATC 17) , Santa Clara, CA, 2017, pp. 181–193
2017
Earlier work this paper cites.
X. Mei, X. Chu, H. Liu, Y. Leung, and Z. Li, “Energy efficient real-time task scheduling on cpu-gpu hybrid clusters,” in IEEE INFOCOM 2017 - IEEE Conference on Computer Communications , May 2017, pp. 1–9
2017
Cited alongside, same era.
V. Chau, X. Chu, H. Liu, and Y.-W. Leung, “Energy efficient job scheduling with dvfs for cpu-gpu heterogeneous systems,” in Proceedings of the Eighth International Conference on Future Energy Systems , ser. e-Energy ’17. ACM, 2017, pp. 1–11
2017
Cited alongside, same era.
Y. You, A. Buluç, and J. Demmel, “Scaling Deep Learning on GPU and Knights Landing Clusters,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’17, New York, NY, USA, 2017, pp. 9:1–9:12
2017
Cited alongside, same era.
M. Cho, U. Finkler, D. Kung, and H. Hunter, “BlueConnect: Decomposing All-Reduce for Deep Learning on Heterogeneous Network Hierarchy,” in the second SysML Conference , Palo Alto, CA, USA, 2019
2019
Later among the works it cites.
S. Shi, Q. Wang, K. Zhao, Z. Tang, Y. Wang, X. Huang, and X. Chu, “A distributed synchronous SGD algorithm with global top-k sparsification for low bandwidth networks,” in 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS) . IEEE, 2019, pp. 2238–2247
2019
Later among the works it cites.
H. Yabuuchi, D. Taniwaki, and S. Omura, “Low-latency job scheduling with preemption for the development of deep learning,” in 2019 USENIX Conference on Operational Machine Learning (OpML 19) , Santa Clara, CA, May 2019, pp. 27–30
2019
Later among the works it cites.
A. Hsu, K. Hu, J. Hung, A. Suresh, and Z. Zhang, “Tony: An orchestrator for distributed machine learning jobs,” in 2019 USENIX Conference on Operational Machine Learning (OpML 19) , Santa Clara, CA, May 2019, pp. 39–41
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
S. Shi, W. Qiang, and X. Chu, “Performance modeling and evaluation of distributed deep learning frameworks on GPUs,” in The 4th International Conference on Big Data Intelligence and Computing (DataCom) . IEEE, 2018, pp. 949–957
2018
Cited alongside, same era.
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally, “Deep gradient compression: Reducing the communication bandwidth for distributed training,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
A. Qiao, A. Aghayev, W. Yu, H. Chen, Q. Ho, G. A. Gibson, and E. P. Xing, “Litz: Elastic framework for high-performance distributed machine learning,” in 2018 USENIX Annual Technical Conference (USENIX ATC 18) , Boston, MA, July 2018, pp. 631–644
2018
Cited alongside, same era.
Y. Peng, Y. Bao, Y. Chen, C. Wu, and C. Guo, “Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters,” in Proceedings of the Thirteenth EuroSys Conference , ser. EuroSys ’18, New York, NY, USA, 2018, pp. 3:1–3:14
2018
Cited alongside, same era.
W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Peng, H. Zhao, Q. Zhang, F. Yang, and L. Zhou, “Gandiva: Introspective Cluster Scheduling for Deep Learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) , Carlsbad, CA, 2018, pp. 595–610
2018
Cited alongside, same era.
Y. Bao, Y. Peng, C. Wu, and Z. Li, “Online job scheduling in distributed machine learning clusters,” in IEEE INFOCOM 2018 - IEEE Conference on Computer Communications , April 2018, pp. 495–503
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Shi, Q. Wang, X. Chu, and B. Li, “A DAG model of synchronous stochastic gradient descent in distributed deep learning,” in 2018 IEEE 24th International Conference on Parallel and Distributed Systems (ICPADS) . IEEE, 2018, pp. 425–432
2018
Cited alongside, same era.
2019
Later among the works it cites.
J. Gu, M. Chowdhury, K. G. Shin, Y. Zhu, M. Jeon, J. Qian, H. Liu, and C. Guo, “Tiresias: A GPU Cluster Manager for Distributed Deep Learning,” in 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) , Boston, MA, 2019, pp. 485–500
2019
Later among the works it cites.
F. Giroire, N. Huin, A. Tomassilli, and S. Pérennes, “When Network Matters: Data Center Scheduling with Network Tasks,” in IEEE INFOCOM 2019 - IEEE Conference on Computer Communications , April 2019, pp. 2278–2286
2019
Later among the works it cites.
C. Chen, X. Ke, T. Zeyl et al. , “Minimum Makespan Workflow Scheduling for Malleable Jobs with Precedence Constraints and Lifetime Resource Demands,” in 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS) , July 2019
2019
Later among the works it cites.
H. Zheng, F. Xu, L. Chen, Z. Zhou, and F. Liu, “Cynthia: Cost-Efficient Cloud Resource Provisioning for Predictable Distributed Deep Neural Network Training,” in 2019 48th International Conference on Parallel Processing (ICPP) , Aug 2019
2019
Later among the works it cites.
Y. Liu, H. Xu, and W. C. Lau, “Online job scheduling with resource packing on a cluster of heterogeneous servers,” in IEEE INFOCOM 2019 - IEEE Conference on Computer Communications , April 2019, pp. 1441–1449
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Bao, Y. Peng, and C. Wu, “Deep Learning-based Job Placement in Distributed Machine Learning Clusters,” in IEEE INFOCOM 2019 - IEEE Conference on Computer Communications , April 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Shi, Q. Wang, X. Chu, B. Li, Y. Qin, R. Liu, and X. Zhao, “Communication-efficient distributed deep learning with merged gradient sparsification on gpus,” in IEEE INFOCOM 2020 - IEEE Conference on Computer Communications , 2020
2020
Closest in time.