Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) like GPT and LLaMA are revolutionizing the AI industry with their sophisticated capabilities.
C. Clos, “A study of non-blocking switching networks,” Bell System Technical Journal , vol. 32, no. 2, pp. 406–424, 1953
1953
Earlier work this paper cites.
L. E. Cannon, A cellular computer to implement the Kalman filter algorithm . Montana State University, 1969
1969
Earlier work this paper cites.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation , vol. 3, no. 1, pp. 79–87, 1991
1991
Earlier work this paper cites.
L. Clarke, I. Glendinning, and R. Hempel, “The mpi message passing interface standard,” in Programming Environments for Massively Parallel Distributed Systems: Working Conference of the IFIP WG 10.3, April 25–29, 1994 . Springer, 1994, pp. 213–218
1994
Earlier work this paper cites.
R. C. Agarwal, S. M. Balle, F. G. Gustavson, M. Joshi, and P. Palkar, “A three-dimensional approach to parallel matrix multiplication,” IBM Journal of Research and Development , vol. 39, no. 5, pp. 575–582, 1995
1995
Earlier work this paper cites.
R. A. Van De Geijn and J. Watts, “Summa: Scalable universal matrix multiplication algorithm,” Concurrency: Practice and Experience , vol. 9, no. 4, pp. 255–274, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , 1998
1998
Earlier work this paper cites.
C. Hopps. (2000) Analysis of an equal-cost multi-path algorithm. [Online]. Available: https://www.rfc-editor.org/rfc/rfc2992.html
2000
Earlier work this paper cites.
G. F. Pfister, “An introduction to the infiniband architecture,” High performance mass storage and parallel I/O , vol. 42, no. 617-632, p. 10, 2001
2001
Earlier work this paper cites.
F. Schmuck and R. Haskin, “ { \{ GPFS } \} : A { \{ Shared-Disk } \} file system for large computing clusters,” in Conference on file and storage technologies (FAST 02) , 2002
2002
Earlier work this paper cites.
W. Gropp, “Mpich2: A new start for mpi implementations,” in Recent Advances in Parallel Virtual Machine and Message Passing Interface: 9th European PVM/MPI Users’ Group Meeting Linz, Austria, September 29–Oktober 2, 2002 Proceedings 9 . Springer, 2002, pp. 7–7
2002
Earlier work this paper cites.
J. Duato, S. Yalamanchili, and L. Ni, Interconnection networks . Morgan Kaufmann, 2003
2003
Earlier work this paper cites.
P. Schwan et al. , “Lustre: Building a file system for 1000-node clusters,” in Proceedings of the 2003 Linux symposium , vol. 2003, 2003, pp. 380–386
2003
Earlier work this paper cites.
E. Gabriel, G. E. Fagg, G. Bosilca, T. Angskun, J. J. Dongarra, J. M. Squyres, V. Sahay, P. Kambadur, B. Barrett, A. Lumsdaine, R. H. Castain, D. J. Daniel, R. L. Graham, and T. S. Woodall, “Open mpi: Goals, concept, and design of a next generation mpi implementation,” in Proceedings, 11th European PVM/MPI Users’ Group Meeting , Budapest, Hungary, September 2004, pp. 97–104
2004
Earlier work this paper cites.
S. Weil, S. A. Brandt, E. L. Miller, D. D. Long, and C. Maltzahn, “Ceph: A scalable, high-performance distributed file system,” in Proceedings of the 7th Conference on Operating Systems Design and Implementation (OSDI’06) , 2006, pp. 307–320
2006
Earlier work this paper cites.
C. Guo, H. Wu, K. Tan, L. Shi, Y. Zhang, and S. Lu, “Dcell: a scalable and fault-tolerant network structure for data centers,” in Proceedings of the ACM SIGCOMM 2008 conference on Data communication , 2008, pp. 75–86
2008
Earlier work this paper cites.
J. Kim, W. J. Dally, S. Scott, and D. Abts, “Technology-driven, highly-scalable dragonfly topology,” ACM SIGARCH Computer Architecture News , vol. 36, no. 3, pp. 77–88, 2008
2008
Earlier work this paper cites.
K. Kong, “Using pci express® as the primary system interconnect in multiroot compute, storage, communications and embedded systems,” White Paper, Integrated Device Technology , p. 12, 2008
2008
Earlier work this paper cites.
C. Guo, G. Lu, D. Li, H. Wu, X. Zhang, Y. Shi, C. Tian, Y. Zhang, and S. Lu, “Bcube: a high performance, server-centric network architecture for modular data centers,” in Proceedings of the ACM SIGCOMM 2009 conference on Data communication , 2009, pp. 63–74
2009
Earlier work this paper cites.
P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,” Journal of Parallel and Distributed Computing , vol. 69, no. 2, pp. 117–124, 2009
2009
Earlier work this paper cites.
Infiniband Trade Association, “Supplement to infiniband architecture specification volume 1 release 1.2.2 annex a16,” pp. 1–17, 2010
2010
Earlier work this paper cites.
——, “Supplement to infiniband architecture specification volume 1 release 1.2.2 annex a17,” pp. 1–17, 2010
2010
Earlier work this paper cites.
K. Shvachko, H. Kuang, S. Radia, and R. Chansler, “The hadoop distributed file system,” in 2010 IEEE 26th symposium on mass storage systems and technologies (MSST) . Ieee, 2010, pp. 1–10
2010
Earlier work this paper cites.
G. Shainer, A. Ayoub, P. Lui, T. Liu, M. Kagan, C. R. Trott, G. Scantlen, and P. S. Crozier, “The development of mellanox/nvidia gpudirect over infiniband—a new model for gpu to gpu communications,” Computer Science-Research and Development , vol. 26, pp. 267–273, 2011
2011
Earlier work this paper cites.
E. Solomonik and J. Demmel, “Communication-optimal parallel 2.5 d matrix multiplication and lu factorization algorithms,” in European Conference on Parallel Processing . Springer, 2011, pp. 90–109
2011
Earlier work this paper cites.
A. Singla, C.-Y. Hong, L. Popa, and P. B. Godfrey, “Jellyfish: Networking data centers randomly,” in 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12) , 2012, pp. 225–238
2012
Earlier work this paper cites.
P. Costa, A. Donnelly, A. Rowstron, and G. O’Shea, “Camdoop: Exploiting in-network aggregation for big data applications,” in 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12) , 2012, pp. 29–42
2012
Earlier work this paper cites.
A. Dixit, P. Prakash, Y. C. Hu, and R. R. Kompella, “On the impact of packet spraying in data center networks,” in 2013 proceedings ieee infocom . IEEE, 2013, pp. 2130–2138
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, and S. Amarasinghe, “Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines,” Acm Sigplan Notices , vol. 48, no. 6, pp. 519–530, 2013
2013
Earlier work this paper cites.
D. K. Panda, K. Tomko, K. Schulz, and A. Majumdar, “The mvapich project: Evolution and sustainability of an open source production quality mpi library for hpc,” in Workshop on Sustainable Software for Science: Practice and Experiences, held in conjunction with Int’l Conference on Supercomputing (WSSPE) , 2013
2013
Earlier work this paper cites.
H. Li, A. Ghodsi, M. Zaharia, S. Shenker, and I. Stoica, “Tachyon: Reliable, memory speed storage for cluster computing frameworks,” in Proceedings of the ACM Symposium on Cloud Computing , 2014, pp. 1–15
2014
Earlier work this paper cites.
L. Mai, L. Rupprecht, A. Alim, P. Costa, M. Migliavacca, P. Pietzuch, and A. L. Wolf, “Netagg: Using middleboxes for application-specific on-path aggregation in data centres,” in Proceedings of the 10th ACM International on Conference on emerging Networking Experiments and Technologies , 2014, pp. 249–262
2014
Earlier work this paper cites.
P. Bosshart, D. Daly, G. Gibb, M. Izzard, N. McKeown, J. Rexford, C. Schlesinger, D. Talayco, A. Vahdat, G. Varghese et al. , “P4: Programming protocol-independent packet processors,” ACM SIGCOMM Computer Communication Review , vol. 44, no. 3, pp. 87–95, 2014
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in 3rd International Conference on Learning Representations , ser. ICLR ’15, 2015
2015
Earlier work this paper cites.
R. Mittal, V. T. Lam, N. Dukkipati, E. Blem, H. Wassel, M. Ghobadi, A. Vahdat, Y. Wang, D. Wetherall, and D. Zats, “Timely: Rtt-based congestion control for the datacenter,” ACM SIGCOMM Computer Communication Review , vol. 45, no. 4, pp. 537–550, 2015
2015
Earlier work this paper cites.
Y. Zhu, H. Eran, D. Firestone, C. Guo, M. Lipshteyn, Y. Liron, J. Padhye, S. Raindel, M. H. Yahia, and M. Zhang, “Congestion control for large-scale rdma deployments,” ACM SIGCOMM Computer Communication Review , vol. 45, no. 4, pp. 523–536, 2015
2015
Earlier work this paper cites.
Y. Zhu, M. Ghobadi, V. Misra, and J. Padhye, “Ecn or delay: Lessons learnt from analysis of dcqcn and timely,” in Proceedings of the 12th International on Conference on emerging Networking EXperiments and Technologies , 2016, pp. 313–327
2016
Earlier work this paper cites.
C. Guo, H. Wu, Z. Deng, G. Soni, J. Ye, J. Padhye, and M. Lipshteyn, “Rdma over commodity ethernet at scale,” in Proceedings of the 2016 ACM SIGCOMM Conference , 2016, pp. 202–215
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: a system for large-scale machine learning,” in 12th USENIX symposium on operating systems design and implementation (OSDI 16) , 2016, pp. 265–283
2016
Earlier work this paper cites.
Nvidia. (2016) Nccl library. [Online]. Available: https://github.com/NVIDIA/nccl
2016
Earlier work this paper cites.
R. L. Graham, D. Bureddy, P. Lui, H. Rosenstock, G. Shainer, G. Bloch, D. Goldenerg, M. Dubman, S. Kotchubievsky, V. Koushnir et al. , “Scalable hierarchical aggregation protocol (sharp): A hardware architecture for efficient data reduction,” in 2016 First International Workshop on Communication Optimizations in HPC (COMHPC) . IEEE, 2016, pp. 1–10
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
NVIDIA, “Nvidia dgx-1 system architecture white paper,” NVIDIA, White Paper, 2017
2017
Earlier work this paper cites.
A. Shpiner, Z. Haramaty, S. Eliad, V. Zdornov, B. Gafni, and E. Zahavi, “Dragonfly+: Low cost topology for scaling datacenters,” in 2017 IEEE 3rd International Workshop on High-Performance Interconnection Networks in the Exascale and Big-Data Era (HiPINEB) . IEEE, 2017, pp. 1–8
2017
Earlier work this paper cites.
“Roce vs. iwarp competitive analysis,” 2017. [Online]. Available: https://network.nvidia.com/sites/default/files/pdf/whitepapers/WP_RoCE_vs_iWARP.pdf
2017
Earlier work this paper cites.
A. Mirhoseini, H. Pham, Q. V. Le, B. Steiner, R. Larsen, Y. Zhou, N. Kumar, M. Norouzi, S. Bengio, and J. Dean, “Device placement optimization with reinforcement learning,” in International conference on machine learning . PMLR, 2017, pp. 2430–2439
2017
Earlier work this paper cites.
G. X. team. (2017) Xla: Optimizing compiler for machine learning. [Online]. Available: https://github.com/openxla/xla
2017
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing, “Poseidon: An efficient communication architecture for distributed deep learning on { \{ GPU } \} clusters,” in 2017 USENIX Annual Technical Conference (USENIX ATC 17) , 2017, pp. 181–193
2017
Earlier work this paper cites.
NVIDIA, “Nvidia data center gpu manager (dcgm),” https://developer.nvidia.com/dcgm, 2017
2017
Earlier work this paper cites.
K. Mandakolathur and S. Jeaugey. (2018) Doubling all2all performance with nvidia collective communication library 2.12. [Online]. Available: https://developer.nvidia.com/blog/nvswitch-leveraging-nvlink-to-maximum-effect/
2018
Earlier work this paper cites.
NVIDIA. (2018) Nvidia dgx-2: The world’s most powerful system for the most complex ai challenges. [Online]. Available: {https://www.nvidia.com/en-us/data-center/dgx-2/}
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Peng, H. Zhao, Q. Zhang et al. , “Gandiva: Introspective cluster scheduling for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) , 2018, pp. 595–610
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young et al. , “Mesh-tensorflow: Deep learning for supercomputers,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
Z. Jia, S. Lin, C. R. Qi, and A. Aiken, “Exploring hidden dimensions in parallelizing convolutional neural networks.” in ICML , 2018, pp. 2279–2288
2018
Earlier work this paper cites.
Y. Gao, L. Chen, and B. Li, “Spotlight: Optimizing device placement for training deep neural networks,” in International Conference on Machine Learning . PMLR, 2018, pp. 1676–1684
2018
Earlier work this paper cites.
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne et al. , “Jax: composable transformations of python+ numpy programs,” 2018
2018
Earlier work this paper cites.
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan et al. , “Ray: A distributed framework for emerging { \{ AI } \} applications,” in 13th USENIX symposium on operating systems design and implementation (OSDI 18) , 2018, pp. 561–577
2018
Earlier work this paper cites.
N. Wang, J. Choi, D. Brand, C.-Y. Chen, and K. Gopalakrishnan, “Training deep neural networks with 8-bit floating point numbers,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
AMD. (2018) Rccl library. [Online]. Available: https://github.com/ROCm/rccl
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Li, R. Miao, H. H. Liu, Y. Zhuang, F. Feng, L. Tang, Z. Cao, M. Zhang, F. Kelly, M. Alizadeh et al. , “Hpcc: High precision congestion control,” in Proceedings of the ACM special interest group on data communication , 2019, pp. 44–58
2019
Earlier work this paper cites.
F. Chowdhury, Y. Zhu, T. Heer, S. Paredes, A. Moody, R. Goldstone, K. Mohror, and W. Yu, “I/o characterization and performance evaluation of beegfs for deep learning,” in Proceedings of the 48th International Conference on Parallel Processing , 2019, pp. 1–10
2019
Earlier work this paper cites.
J. Gu, M. Chowdhury, K. G. Shin, Y. Zhu, M. Jeon, J. Qian, H. Liu, and C. Guo, “Tiresias: A { \{ GPU } \} cluster manager for distributed deep learning,” in 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) , 2019, pp. 485–500
2019
Earlier work this paper cites.
A. Li, S. L. Song, J. Chen, J. Li, X. Liu, N. R. Tallent, and K. J. Barker, “Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect,” IEEE Transactions on Parallel and Distributed Systems , vol. 31, no. 1, pp. 94–110, 2019
2019
Earlier work this paper cites.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM symposium on operating systems principles , 2019, pp. 1–15
2019
Earlier work this paper cites.
M. Jeon, S. Venkataraman, A. Phanishayee, J. Qian, W. Xiao, and F. Yang, “Analysis of large-scale multi-tenant GPU clusters for DNN training workloads,” in 2019 USENIX Annual Technical Conference , ser. USENIX ATC ’19. USENIX Association, 2019, pp. 947–960
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
Z. Jia, M. Zaharia, and A. Aiken, “Beyond data and model parallelism for deep neural networks.” Proceedings of Machine Learning and Systems , vol. 1, pp. 1–13, 2019
2019
Earlier work this paper cites.
M. Wang, C.-c. Huang, and J. Li, “Supporting very large models using automatic dataflow graph partitioning,” in Proceedings of the Fourteenth EuroSys Conference 2019 , 2019, pp. 1–17
2019
Earlier work this paper cites.
L. Song, J. Mao, Y. Zhuo, X. Qian, H. Li, and Y. Chen, “Hypar: Towards hybrid parallelism for deep learning accelerator array,” in 2019 IEEE international symposium on high performance computer architecture (HPCA) . IEEE, 2019, pp. 56–68
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
P. Tillet, H.-T. Kung, and D. Cox, “Triton: an intermediate language and compiler for tiled neural network computations,” in Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages , 2019, pp. 10–19
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
X. Sun, J. Choi, C.-Y. Chen, N. Wang, S. Venkataramani, V. V. Srinivasan, X. Cui, W. Zhang, and K. Gopalakrishnan, “Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
Z. Jia, O. Padon, J. Thomas, T. Warszawski, M. Zaharia, and A. Aiken, “Taso: optimizing deep learning computation with automatic generation of graph substitutions,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 47–62
2019
Earlier work this paper cites.
K. Yang, Y.-F. Chen, G. Roumpos, C. Colby, and J. Anderson, “High performance monte carlo simulation of ising model on tpu clusters,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2019, pp. 1–15
2019
Earlier work this paper cites.
S. Jeaugey. (2019) Massively scale your deep learning training with nccl 2.4. [Online]. Available: https://developer.nvidia.com/blog/massively-scale-deep-learning-training-nccl-2-4/
2019
Earlier work this paper cites.
M. Cho, U. Finkler, D. Kung, and H. Hunter, “Blueconnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy,” Proceedings of Machine Learning and Systems , vol. 1, pp. 241–251, 2019
2019
Earlier work this paper cites.
P. Sun, Y. Wen, R. Han, W. Feng, and S. Yan, “Gradientflow: Optimizing network performance for large-scale distributed dnn training,” IEEE Transactions on Big Data , vol. 8, no. 2, pp. 495–507, 2019
2019
Earlier work this paper cites.
A. Jayarajan, J. Wei, G. Gibson, A. Fedorova, and G. Pekhimenko, “Priority-based parameter propagation for distributed dnn training,” Proceedings of Machine Learning and Systems , vol. 1, pp. 132–145, 2019
2019
Earlier work this paper cites.
S. H. Hashemi, S. Abdu Jyothi, and R. Campbell, “Tictac: Accelerating distributed deep learning with communication scheduling,” Proceedings of Machine Learning and Systems , vol. 1, pp. 418–430, 2019
2019
Earlier work this paper cites.
Y. Peng, Y. Zhu, Y. Chen, Y. Bao, B. Yi, C. Lan, C. Wu, and C. Guo, “A generic communication scheduler for distributed dnn training acceleration,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 16–29
2019
Earlier work this paper cites.
F. Yang, Z. Wang, X. Ma, G. Yuan, and X. An, “Switchagg: A further step towards in-network computation,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2019, pp. 185–185
2019
Earlier work this paper cites.
Y. Li, I.-J. Liu, Y. Yuan, D. Chen, A. Schwing, and J. Huang, “Accelerating distributed reinforcement learning with in-switch computing,” in Proceedings of the 46th International Symposium on Computer Architecture , 2019, pp. 279–291
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 3505–3506
2020
Earlier work this paper cites.
R. Mayer and H.-A. Jacobsen, “Scalable deep learning on distributed infrastructures: Challenges, techniques, and tools,” ACM Computing Surveys (CSUR) , vol. 53, no. 1, pp. 1–37, 2020
2020
Earlier work this paper cites.
S. Naffziger, K. Lepak, M. Paraschou, and M. Subramony, “2.2 amd chiplet architecture for high-performance server and desktop products,” in 2020 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2020, pp. 44–45
2020
Earlier work this paper cites.
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson, “A domain-specific supercomputer for training deep neural networks,” Communications of the ACM , vol. 63, no. 7, pp. 67–78, 2020
2020
Earlier work this paper cites.
T. Norrie, N. Patil, D. H. Yoon, G. Kurian, S. Li, J. Laudon, C. Young, N. P. Jouppi, and D. A. Patterson, “Google’s training chips revealed: Tpuv2 and tpuv3.” in Hot Chips Symposium , 2020, pp. 1–70
2020
Earlier work this paper cites.
G. Kumar, N. Dukkipati, K. Jang, H. M. Wassel, X. Wu, B. Montazeri, Y. Wang, K. Springborn, C. Alfeld, M. Ryan et al. , “Swift: Delay is simple and effective for congestion control in the datacenter,” in Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, technologies, architectures, and protocols for computer communication , 2020, pp. 514–528
2020
Earlier work this paper cites.
P. Taheri, D. Menikkumbura, E. Vanini, S. Fahmy, P. Eugster, and T. Edsall, “Rocc: robust congestion control for rdma,” in Proceedings of the 16th International conference on emerging networking experiments and technologies , 2020, pp. 17–30
2020
Earlier work this paper cites.
A. V. Kumar and M. Sivathanu, “Quiver: An informed storage cache for deep learning,” in 18th USENIX Conference on File and Storage Technologies (FAST 20) , 2020, pp. 283–296
2020
Earlier work this paper cites.
K. Mahajan, A. Balasubramanian, A. Singhvi, S. Venkataraman, A. Akella, A. Phanishayee, and S. Chawla, “Themis: Fair and efficient { \{ GPU } \} cluster scheduling,” in 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20) , 2020, pp. 289–304
2020
Earlier work this paper cites.
D. Narayanan, K. Santhanam, F. Kazhamiaka, A. Phanishayee, and M. Zaharia, “Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads,” in 14th USENIX Symposium on Operating Systems Design and Implementation , ser. OSDI ’20. USENIX Association, 2020, pp. 481–498
2020
Earlier work this paper cites.
S. Chaudhary, R. Ramjee, M. Sivathanu, N. Kwatra, and S. Viswanatha, “Balancing efficiency and fairness in heterogeneous gpu clusters for deep learning,” in Proceedings of the Fifteenth European Conference on Computer Systems , ser. EuroSys ’20. Association for Computing Machinery, 2020
2020
Earlier work this paper cites.
Z. Zhang, C. Chang, H. Lin, Y. Wang, R. Arora, and X. Jin, “Is network the bottleneck of distributed training?” in Proceedings of the Workshop on Network Meets AI & ML , 2020, pp. 8–13
2020
Earlier work this paper cites.
“Infiniband product guide,” 2020, accessed: 2024-07-01. [Online]. Available: https://network.nvidia.com/files/doc-2020/br-infiniband-product-guide.pdf
2020
Earlier work this paper cites.
J. Dong, Z. Cao, T. Zhang, J. Ye, S. Wang, F. Feng, L. Zhao, X. Liu, L. Song, L. Peng et al. , “Eflops: Algorithm and system co-design for a high performance distributed training platform,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 610–622
2020
Earlier work this paper cites.
W. Xiao, S. Ren, Y. Li, Y. Zhang, P. Hou, Z. Li, Y. Feng, W. Lin, and Y. Jia, “ { \{ AntMan } \} : Dynamic scaling on { \{ GPU } \} clusters for deep learning,” in 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) , 2020, pp. 533–548
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Earlier work this paper cites.
J. M. Tarnawski, A. Phanishayee, N. Devanur, D. Mahajan, and F. Nina Paravecino, “Efficient algorithms for device placement of dnn graph operators,” Advances in Neural Information Processing Systems , vol. 33, pp. 15 451–15 463, 2020
2020
Earlier work this paper cites.
J. H. Park, G. Yun, M. Y. Chang, N. T. Nguyen, S. Lee, J. Choi, S. H. Noh, and Y.-r. Choi, “Hetpipe: Enabling large dnn training on (whimpy) heterogeneous gpu clusters through integration of pipelined model parallelism and data parallelism,” in 2020 USENIX Annual Technical Conference (USENIX ATC 20) , 2020, pp. 307–321
2020
Earlier work this paper cites.
L. Song, F. Chen, Y. Zhuo, X. Qian, H. Li, and Y. Chen, “Accpar: Tensor partitioning for heterogeneous deep learning accelerators,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 342–355
2020
Earlier work this paper cites.
L. Zheng, C. Jia, M. Sun, Z. Wu, C. H. Yu, A. Haj-Ali, Y. Wang, J. Yang, D. Zhuo, K. Sen et al. , “Ansor: Generating { \{ High-Performance } \} tensor programs for deep learning,” in 14th USENIX symposium on operating systems design and implementation (OSDI 20) , 2020, pp. 863–879
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. Jain, A. Jain, A. Nrusimha, A. Gholami, P. Abbeel, J. Gonzalez, K. Keutzer, and I. Stoica, “Checkmate: Breaking the memory wall with optimal tensor rematerialization,” Proceedings of Machine Learning and Systems , vol. 2, pp. 497–511, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
L. Luo, P. West, A. Krishnamurthy, L. Ceze, and J. Nelson, “Plink: Discovering and exploiting datacenter network locality for efficient cloud-based distributed training,” Proc. of MLSys , 2020
2020
Earlier work this paper cites.
G. Wang, S. Venkataraman, A. Phanishayee, N. Devanur, J. Thelin, and I. Stoica, “Blink: Fast and generic collectives for distributed ml,” pp. 172–186, 2020
2020
Earlier work this paper cites.
Y. Bao, Y. Peng, Y. Chen, and C. Wu, “Preemptive all-reduce scheduling for expediting distributed dnn training,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications . IEEE, 2020, pp. 626–635
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
B. Nicolae, J. Li, J. M. Wozniak, G. Bosilca, M. Dorier, and F. Cappello, “Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models,” in 2020 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID) . IEEE, 2020, pp. 172–181
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Choquette, E. Lee, R. Krashinsky, V. Balan, and B. Khailany, “3.2 the a100 datacenter gpu and ampere architecture,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IEEE, 2021, pp. 48–50
2021
Earlier work this paper cites.
M. Khani, M. Ghobadi, M. Alizadeh, Z. Zhu, M. Glick, K. Bergman, A. Vahdat, B. Klenk, and E. Ebrahimi, “Sip-ml: high-bandwidth optical network interconnects for machine learning training,” in Proceedings of the 2021 ACM SIGCOMM 2021 Conference , 2021, pp. 657–675
2021
Cited alongside, same era.
S. Pan, T. Stavrinos, Y. Zhang, A. Sikaria, P. Zakharov, A. Sharma, M. Shuey, R. Wareing, M. Gangapuram, G. Cao et al. , “Facebook’s tectonic filesystem: Efficiency from exascale,” in 19th USENIX Conference on File and Storage Technologies (FAST 21) , 2021, pp. 217–231
2021
Cited alongside, same era.
A. Qiao, S. K. Choe, S. J. Subramanya, W. Neiswanger, Q. Ho, H. Zhang, G. R. Ganger, and E. P. Xing, “Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning,” in 15th USENIX Symposium on Operating Systems Design and Implementation , ser. OSDI ’21. USENIX Association, 2021, pp. 1–18
2021
Cited alongside, same era.
Z. Lai, S. Li, X. Tang, K. Ge, W. Liu, Y. Duan, L. Qiao, and D. Li, “Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,” IEEE Transactions on Parallel and Distributed Systems , vol. 34, no. 5, pp. 1466–1478, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Miao, Y. Shi, Z. Yang, B. Cui, and Z. Jia, “Sdpipe: A semi-decentralized framework for heterogeneity-aware pipeline-parallel training,” Proceedings of the VLDB Endowment , vol. 16, no. 9, pp. 2354–2363, 2023
2023
Later among the works it cites.
J. Zhang, G. Niu, Q. Dai, H. Li, Z. Wu, F. Dong, and Z. Wu, “Pipepar: Enabling fast dnn pipeline parallel training in heterogeneous gpu clusters,” Neurocomputing , vol. 555, p. 126661, 2023
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Blöcher, L. Wang, P. Eugster, and M. Schmidt, “Switches for hire: Resource scheduling for data center in-network computing,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2021, pp. 268–285
2021
Cited alongside, same era.
S. Li and T. Hoefler, “Chimera: efficiently training large-scale neural networks with bidirectional pipelines,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–14
2021
Cited alongside, same era.
Q. Hu, P. Sun, S. Yan, Y. Wen, and T. Zhang, “Characterization and prediction of deep learning workloads in large-scale gpu datacenters,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’21, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. Fan, Y. Rong, C. Meng, Z. Cao, S. Wang, Z. Zheng, C. Wu, G. Long, J. Yang, L. Xia et al. , “Dapple: A pipelined data parallel approach for training large models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 431–445
2021
Cited alongside, same era.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–15
2021
Cited alongside, same era.
Z. Li, S. Zhuang, S. Guo, D. Zhuo, H. Zhang, D. Song, and I. Stoica, “Terapipe: Token-level pipeline parallelism for training large-scale language models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6543–6552
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Later among the works it cites.
M. Ryabinin, T. Dettmers, M. Diskin, and A. Borzunov, “Swarm parallelism: Training large models can be surprisingly communication-efficient,” in International Conference on Machine Learning . PMLR, 2023, pp. 29 416–29 440
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4195–4205
2023
Later among the works it cites.
S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y. He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,” in Proceedings of the 37th International Conference on Supercomputing , 2023, pp. 203–214
2023
Later among the works it cites.
S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y. Liu, B. Wang, and Y. You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing , 2023, pp. 766–775
2023
Later among the works it cites.
J. Wang, Y. Lu, B. Yuan, B. Chen, P. Liang, C. De Sa, C. Re, and C. Zhang, “Cocktailsgd: Fine-tuning foundation models over 500mbps networks,” in International Conference on Machine Learning . PMLR, 2023, pp. 36 058–36 076
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhai, C. Jiang, L. Wang, X. Jia, S. Zhang, Z. Chen, X. Liu, and Y. Zhu, “Bytetransformer: A high-performance transformer boosted for variable-length inputs,” in 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2023, pp. 344–355
2023
Later among the works it cites.
G. Huang, Y. Bai, L. Liu, Y. Wang, B. Yu, Y. Ding, and Y. Xie, “Alcop: Automatic load-compute pipelining in deep learning compiler for ai-gpus,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Later among the works it cites.
S. Zheng, S. Chen, P. Song, R. Chen, X. Li, S. Yan, D. Lin, J. Leng, and Y. Liang, “Chimera: An analytical optimizing framework for effective compute-intensive operators fusion,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 1113–1126
2023
Later among the works it cites.
Y. Shi, Z. Yang, J. Xue, L. Ma, Y. Xia, Z. Miao, Y. Guo, F. Yang, and L. Zhou, “Welder: Scheduling deep learning memory access via tile-graph,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) , 2023, pp. 701–718
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Xi, C. Li, J. Chen, and J. Zhu, “Training transformers with 4-bit integers,” Advances in Neural Information Processing Systems , vol. 36, pp. 49 146–49 168, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2818–2829
2023
Later among the works it cites.
J. Choquette, “Nvidia hopper h100 gpu: Scaling performance,” IEEE Micro , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
PyTorch. (2023) Pytorch expandable segments. [Online]. Available: https://github.com/pytorch/pytorch/pull/96995
2023
Later among the works it cites.
Y. Feng, M. Xie, Z. Tian, S. Wang, Y. Lu, and J. Shu, “Mobius: Fine tuning large-scale models on commodity gpu servers,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 489–501
2023
Later among the works it cites.
S.-F. Lin, Y.-J. Chen, H.-Y. Cheng, and C.-L. Yang, “Tensor movement orchestration in multi-gpu training systems,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 1140–1152
2023
Later among the works it cites.
X. Nie, Y. Liu, F. Fu, J. Xue, D. Jiao, X. Miao, Y. Tao, and B. Cui, “Angel-ptm: A scalable and economical large-scale pre-training system in tencent,” Proceedings of the VLDB Endowment , vol. 16, no. 12, pp. 3781–3794, 2023
2023
Later among the works it cites.
A. Shah, V. Chidambaram, M. Cowan, S. Maleki, M. Musuvathi, T. Mytkowicz, J. Nelson, O. Saarikivi, and R. Singh, “ { \{ TACCL } \} : Guiding collective algorithm synthesis using communication sketches,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , 2023, pp. 593–612
2023
Later among the works it cites.
F. Li, S. Zhao, Y. Qing, X. Chen, X. Guan, S. Wang, G. Zhang, and H. Cui, “Fold3d: Rethinking and parallelizing computational and communicational tasks in the training of large dnn models,” IEEE Transactions on Parallel and Distributed Systems , vol. 34, no. 5, pp. 1432–1449, 2023
2023
Later among the works it cites.
K. Mahajan, C.-H. Chu, S. Sridharan, and A. Akella, “Better together: Jointly optimizing { \{ ML } \} collective scheduling and execution planning using { \{ SYNDICATE } \} ,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , 2023, pp. 809–824
2023
Later among the works it cites.
L. Zhang, S. Shi, X. Chu, W. Wang, B. Li, and C. Liu, “Dear: Accelerating distributed deep learning with fine-grained all-reduce pipelining,” in 2023 IEEE 43rd International Conference on Distributed Computing Systems (ICDCS) . IEEE, 2023, pp. 142–153
2023
Later among the works it cites.
S. Li, Z. Lai, Y. Hao, W. Liu, K. Ge, X. Deng, D. Li, and K. Lu, “Automated tensor model parallelism with overlapped communication for efficient foundation model training,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Chen, Y. Hua, R. Bai, and J. Huang, “A cost-efficient failure-tolerant scheme for distributed dnn training,” in 2023 IEEE 41st International Conference on Computer Design (ICCD) . IEEE, 2023, pp. 150–157
2023
Later among the works it cites.
Z. Wang, Z. Jia, S. Zheng, Z. Zhang, X. Fu, T. E. Ng, and Y. Wang, “Gemini: Fast failure recovery in distributed training with in-memory checkpoints,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 364–381
2023
Later among the works it cites.
2023
Later among the works it cites.
I. Jang, Z. Yang, Z. Zhang, X. Jin, and M. Chowdhury, “Oobleck: Resilient distributed training of large models using pipeline templates,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 382–395
2023
Later among the works it cites.
J. Thorpe, P. Zhao, J. Eyolfson, Y. Qiao, Z. Jia, M. Zhang, R. Netravali, and G. H. Xu, “Bamboo: Making preemptible instances resilient for affordable training of large { \{ DNNs } \} ,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , 2023, pp. 497–513
2023
Later among the works it cites.
Y. Gao, X. Shi, H. Lin, H. Zhang, H. Wu, R. Li, and M. Yang, “An empirical study on quality issues of deep learning platform,” in 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) . IEEE, 2023, pp. 455–466
2023
Later among the works it cites.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Later among the works it cites.
M. Steinman. (2023) Taking stock of new data center computing paradigm. [Online]. Available: https://www.eetimes.com/taking-stock-of-new-data-center-computing-paradigm/
2023
Later among the works it cites.
LlamaTeam. (2024) The llama 3 herd of models. [Online]. Available: https://ai.meta.com/research/publications/the-llama-3-herd-of-models
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Sun, X. Zhang, Z. He, P. Li, Q. Cheng, X. Liu, H. Yan, Y. Shao, Q. Tang, S. Zhang et al. , “Moss: An open conversational large language model,” Machine Intelligence Research , pp. 1–18, 2024
2024
Closest in time.
Q. Hu, Z. Ye, Z. Wang, G. Wang, M. Zhang, Q. Chen, P. Sun, D. Lin, X. Wang, Y. Luo et al. , “Characterization of large language model development in the datacenter,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 709–729
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Qian, Y. Xi, J. Cao, J. Gao, Y. Xu, Y. Guan, B. Fu, X. Shi, F. Zhu, R. Miao et al. , “Alibaba hpn: A data center network for large language model training,” in Proceedings of the ACM SIGCOMM 2024 Conference , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Rajasekaran, M. Ghobadi, and A. Akella, “ { \{ CASSINI } \} : { \{ Network-Aware } \} job scheduling in machine learning clusters,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 1403–1420
2024
Closest in time.
H. Wang, H. Tian, J. Chen, X. Wan, J. Xia, G. Zeng, W. Bai, J. Jiang, Y. Wang, and K. Chen, “Towards { \{ Domain-Specific } \} network transport for distributed { \{ DNN } \} training,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 1421–1443
2024
Closest in time.
2024
Closest in time.
S. Dash, I. R. Lyngaas, J. Yin, X. Wang, R. Egele, J. A. Ellis, M. Maiterth, G. Cong, F. Wang, and P. Balaprakash, “Optimizing distributed training on frontier for large language models,” in ISC High Performance 2024 Research Paper Proceedings (39th International Conference) . Prometeus GmbH, 2024, pp. 1–11
2024
Closest in time.
L. Dai, H. Qi, W. Chen, and X. Lu, “High-speed data communication with advanced networks in large language model training,” IEEE Micro , 2024
2024
Closest in time.
Nvidia. (2024) Dgx superpod architecture. [Online]. Available: https://docs.nvidia.com/dgx-superpod/reference-architecture-scalable-infrastructure-h100/latest/dgx-superpod-architecture.html
2024
Closest in time.
Kevin Lee, Adi Gangidi and Mathew Oldham. Building Meta’s GenAI Infrastructure. [Online]. Available: {https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure}
2024
Closest in time.
2024
Closest in time.
Z. Ye, W. Gao, Q. Hu, P. Sun, X. Wang, Y. Luo, T. Zhang, and Y. Wen, “Deep learning workload scheduling in gpu datacenters: A survey,” ACM Computing Surveys , vol. 56, p. 1–38, 2024
2024
Closest in time.
Z. Lin, Y. Miao, G. Xu, C. Li, O. Saarikivi, S. Maleki, and F. Yang, “Tessel: Boosting distributed execution of large dnn models via flexible schedule search,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2024, pp. 803–816
2024
Closest in time.
2024
Closest in time.
C. Jiang, Z. Jia, S. Zheng, Y. Wang, and C. Wu, “Dynapipe: Optimizing multi-task training through dynamic pipelines,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 542–559
2024
Closest in time.
J. Huang, Z. Zhang, S. Zheng, F. Qin, and Y. Wang, “Distmm: Accelerating distributed multimodal model training,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 1157–1171
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Sun, H. Cao, Y. Wang, G. Feng, S. Chen, H. Wang, and W. Chen, “Adapipe: Optimizing pipeline parallelism with adaptive recomputation and partitioning,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 86–100
2024
Closest in time.
2024
Closest in time.
S. A. Jacobs, M. Tanaka, C. Zhang, M. Zhang, R. Y. Aminadabi, S. L. Song, S. Rajbhandari, and Y. He, “System optimizations for enabling training of extreme long sequence transformer models,” in Proceedings of the 43rd ACM Symposium on Principles of Distributed Computing , 2024, pp. 121–130
2024
Closest in time.
2024
Closest in time.
H. Liu, M. Zaharia, and P. Abbeel, “Ring attention with blockwise transformers for near-infinite context,” International Conference on Learning Representations(ICLR) , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Shi, X. Pan, Q. Wang, C. Liu, X. Ren, Z. Hu, Y. Yang, B. Li, and X. Chu, “Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 236–249
2024
Closest in time.
H. Chen, C. H. Yu, S. Zheng, Z. Zhang, Z. Zhang, and Y. Wang, “Slapo: A schedule language for progressive optimization of large deep learning model training,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024, pp. 1095–1111
2024
Closest in time.
G. Liu, Y. Miao, Z. Lin, X. Shi, S. Maleki, F. Yang, Y. Bao, and S. Wang, “Aceso: Efficient parallel dnn training through iterative bottleneck alleviation,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 163–181
2024
Closest in time.
2024
Closest in time.
J. Chen, S. Li, R. Guo, J. Yuan, and T. Hoefler, “Autoddl: Automatic distributed deep learning with near-optimal bandwidth cost,” IEEE Transactions on Parallel and Distributed Systems , 2024
2024
Closest in time.
Y. Wang, Y. Jiang, X. Miao, F. Fu, S. Zhu, X. Nie, Y. Tu, and B. Cui, “Improving automatic parallel training via balanced memory workload optimization,” IEEE Transactions on Knowledge and Data Engineering , 2024
2024
Closest in time.
S. Zhang, L. Diao, C. Wu, Z. Cao, S. Wang, and W. Lin, “Hap: Spmd dnn training on heterogeneous gpu clusters with automated program synthesis,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 524–541
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Lei, Y. Jin, M. Zhai, K. Huang, H. Ye, and J. Zhai, “ { \{ PUZZLE } \} : Efficiently aligning large language models through { \{ Light-Weight } \} context switch,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24) , 2024, pp. 127–140
2024
Closest in time.
F. Strati, P. Elvinger, T. Kerimoglu, and A. Klimovic, “Ml training with cloud gpu shortages: Is cross-region the answer?” in Proceedings of the 4th Workshop on Machine Learning and Systems , 2024, pp. 107–116
2024
Closest in time.
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
H. Liu and P. Abbeel, “Blockwise parallel transformers for large context models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
R. Wu, X. Zhu, J. Chen, S. Liu, T. Zheng, X. Liu, and H. An, “Swattention: designing fast and memory-efficient attention for a new sunway supercomputer,” The Journal of Supercomputing , pp. 1–24, 2024
2024
Closest in time.
J. Ansel, E. Yang, H. He, N. Gimelshein, A. Jain, M. Voznesensky, B. Bao, P. Bell, D. Berard, E. Burovski et al. , “Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024, pp. 929–947
2024
Closest in time.
M. Ibrahim, S. Aga, A. Li, S. Pati, and M. Islam, “Jit-q: Just-in-time quantization with processing-in-memory for efficient ml training,” Proceedings of Machine Learning and Systems , vol. 6, pp. 46–59, 2024
2024
Closest in time.
M. Li, R. B. Basat, S. Vargaftik, C. Lao, K. Xu, M. Mitzenmacher, and M. Yu, “ { \{ THC } \} : Accelerating distributed deep learning using tensor homomorphic compression,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 1191–1211
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Zhang, S. Ma, P. Liu, and J. Yuan, “Coop: Memory is not a commodity,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
T. Yuan, Y. Liu, X. Ye, S. Zhang, J. Tan, B. Chen, C. Song, and D. Zhang, “Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism,” 2024 USENIX Annual Technical Conference (USENIX ATC 24) , 2024
2024
Closest in time.
Q. Chen, Q. Hu, G. Wang, Y. Xiong, T. Huang, X. Chen, Y. Gao, H. Yan, Y. Wen, T. Zhang, and P. Sun, “Amsp: Reducing communication overhead of zero for efficient llm training,” 2024
2024
Closest in time.
A. Imanishi and Z. Xu, “A Heuristic for Periodic Memory Allocation with Little Fragmentation to Train Neural Networks,” in Proceedings of the 2024 ACM SIGPLAN International Symposium on Memory Management . ACM, 2024, pp. 82–94
2024
Closest in time.
C. Guo, R. Zhang, J. Xu, J. Leng, Z. Liu, Z. Huang, M. Guo, H. Wu, S. Zhao, J. Zhao et al. , “Gmlake: Efficient and transparent gpu memory defragmentation for large-scale dnn training with virtual memory stitching,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024, pp. 450–466
2024
Closest in time.
H. Jang, J. Song, J. Jung, J. Park, Y. Kim, and J. Lee, “Smart-infinity: Fast large language model training using near-storage processing on a real system,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2024, pp. 345–360
2024
Closest in time.
2024
Closest in time.
D. Yu, L. Shen, H. Hao, W. Gong, H. Wu, J. Bian, L. Dai, and H. Xiong, “Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,” IEEE Transactions on Services Computing , 2024
2024
Closest in time.
Z. Zhang, Y. Xia, H. Wang, D. Yang, C. Hu, X. Zhou, and D. Cheng, “Mpmoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,” IEEE Transactions on Parallel and Distributed Systems , 2024
2024
Closest in time.
S. Li, K. Lu, Z. Lai, W. Liu, K. Ge, and D. Li, “A multidimensional communication scheduling method for hybrid parallel dnn training,” IEEE Transactions on Parallel and Distributed Systems , 2024
2024
Closest in time.
C. Chen, X. Li, Q. Zhu, J. Duan, P. Sun, X. Zhang, and C. Yang, “Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 178–191
2024
Closest in time.
S. Pati, S. Aga, M. Islam, N. Jayasena, and M. D. Sinclair, “T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024, pp. 1146–1164
2024
Closest in time.
P. Chen, W. Zhang, S. He, Y. Gu, Z. Peng, K. Huang, X. Zhan, W. Chen, Y. Zheng, Z. Wang et al. , “Optimizing large model training through overlapped activation recomputation,” 2024
2024
Closest in time.
H. Zhu, W. Jiang, Q. Hong, and Z. Guo, “When in-network computing meets distributed machine learning,” IEEE Network , 2024
2024
Closest in time.
Y. Zu, A. Ghaffarkhah, H.-V. Dang, B. Towles, S. Hand, S. Huda, A. Bello, A. Kolbasov, A. Rezaei, D. Du, S. Lacy, H. Wang, A. Wisner, C. Lewis, and H. Bahini, “Resiliency at Scale: Managing Google’s TPUv4 Machine Learning Supercomputer,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 761–774
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Xiong, Y. Jiang, Z. Yang, L. Qu, G. Zhao, S. Liu, D. Zhong, B. Pinzur, J. Zhang, Y. Wang et al. , “Superbench: Improving cloud ai infrastructure reliability with proactive validation,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24) , 2024, pp. 835–850
2024
Closest in time.
T. Gupta, S. Krishnan, R. Kumar, A. Vijeev, B. Gulavani, N. Kwatra, R. Ramjee, and M. Sivathanu, “Just-in-time checkpointing: Low cost error recovery from deep learning training failures,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 1110–1125
2024
Closest in time.
A. Group. (2024) Dlrover: An automatic distributed deep learning system. [Online]. Available: https://github.com/intelligent-machine-learning/dlrover
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Xu, T. Zhou, M. Ma, C. Deng, Q. Dai, and L. Fang, “Large-scale photonic chiplet taichi empowers 160-tops/w artificial general intelligence,” Science , vol. 384, no. 6692, pp. 202–209, 2024
2024
Closest in time.