L. Schrage, “A proof of the optimality of the shortest remaining processing time discipline,” Operations Research
1968
Earlier work this paper cites.
C.-Y. Hong, M. Caesar, and P. B. Godfrey, “Finishing flows quickly with preemptive scheduling,” in ACM SIGCOMM
2012
Earlier work this paper cites.
M. Alizadeh, S. Yang, M. Sharif, S. Katti, N. McKeown, B. Prabhakar, and S. Shenker, “pfabric: Minimal near-optimal datacenter transport,” SIGCOMM CCR
2013
Earlier work this paper cites.
M. Chowdhury, Y. Zhong, and I. Stoica, “Efficient coflow scheduling with varys,” in ACM SIGCOMM
2014
Earlier work this paper cites.
W. Bai, L. Chen, K. Chen, D. Han, C. Tian, and H. Wang, “Information-agnostic flow scheduling for commodity data centers,” in USENIX OSDI
2015
Earlier work this paper cites.
M. Chowdhury and I. Stoica, “Efficient coflow scheduling without prior knowledge,” SIGCOMM CCR
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Neural Information Processing Systems
2017
Earlier work this paper cites.
C. Olston, N. Fiedel, K. Gorovoy, J. Harmsen, L. Lao, F. Li, V. Rajashekhar, S. Ramesh, and J. Soyke, “Tensorflow-serving: Flexible, high-performance ml serving,” arXiv
2017
Earlier work this paper cites.
G. Prekas, M. Kogias, and E. Bugnion, “Zygos: Achieving low tail latency for microsecond-scale networked tasks,” in ACM SOSP
2017
Earlier work this paper cites.
D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica, “Clipper: A low-latency online prediction serving system.,” in USENIX NSDI
2017
Earlier work this paper cites.
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica, “Ray: A distributed framework for emerging AI applications,” in USENIX OSDI
2018
Earlier work this paper cites.
K. Kaffes, T. Chong, J. T. Humphries, A. Belay, D. Mazières, and C. Kozyrakis, “Shinjuku: Preemptive scheduling for μ \mu second-scale tail latency,” in USENIX NSDI
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Neural Information Processing Systems
2019
Earlier work this paper cites.
N. Corporation, “Triton inference server: An optimized cloud and edge inferencing solution.,” 2019
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” arXiv
2019
Earlier work this paper cites.
N. Corporation, “Fastertransformer,” 2019
2019
Earlier work this paper cites.
N. Shazeer, “Fast transformer decoding: One write-head is all you need,” arXiv
2019
Earlier work this paper cites.
J. Gu, M. Chowdhury, K. G. Shin, Y. Zhu, M. Jeon, J. Qian, H. H. Liu, and C. Guo, “Tiresias: A gpu cluster manager for distributed deep learning.,” in USENIX NSDI
2019
Earlier work this paper cites.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: Generalized pipeline parallelism for dnn training,” in ACM SOSP
2019
Earlier work this paper cites.