Fetching the paper…
Reading the bibliography…
Reliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to grow.
1909
Earlier work this paper cites.
J. W. Young, “A first order approximation to the optimum checkpoint interval,” Commun. ACM , vol. 17, pp. 530–531, 1974. [Online]. Available: https://doi.org/10.1145/361147.361115
1974
Earlier work this paper cites.
L. G. Valiant, “A bridging model for parallel computation,” Commun. ACM , vol. 33, no. 8, p. 103–111, Aug. 1990. [Online]. Available: https://doi.org/10.1145/79173.79181
1990
Earlier work this paper cites.
M. Harchol-Balter, K. Sigman, and A. Wierman, “Asymptotic convergence of scheduling policies with respect to slowdown,” Performance Evaluation , vol. 49, no. 1, pp. 241–256, 2002, performance 2002. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0166531602001323
2002
Earlier work this paper cites.
F. Hansen and G. K. Pedersen, “Jensen’s operator inequality,” Bulletin of the London Mathematical Society , vol. 35, no. 04, p. 553–564, 2003
2003
Earlier work this paper cites.
A. B. Yoo, M. A. Jette, and M. Grondona, “Slurm: Simple linux utility for resource management,” in Workshop on job scheduling strategies for parallel processing . Springer, 2003, pp. 44–60
2003
Earlier work this paper cites.
J. Daly, “A higher order estimate of the optimum checkpoint interval for restart dumps,” Future Generation Computer Systems , vol. 22, no. 3, pp. 303–312, 2006. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167739X04002213
2006
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild!: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in Neural Information Processing Systems , vol. 24, 2011. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2011/file/218a0aefd1d1a4be65601cc6ddc1520e-Paper.pdf
2011
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. a. Ranzato, A. Senior, P. Tucker, K. Yang, Q. Le, and A. Ng, “Large scale distributed deep networks,” in Advances in Neural Information Processing Systems , vol. 25, 2012. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2012/file/6aca97005c68f1206823815f66102863-Paper.pdf
2012
Earlier work this paper cites.
C. Reiss, A. Tumanov, G. R. Ganger, R. H. Katz, and M. A. Kozuch, “Heterogeneity and dynamicity of clouds at scale: Google trace analysis,” in Proceedings of the Third ACM Symposium on Cloud Computing , ser. SoCC ’12. New York, NY, USA: Association for Computing Machinery, 2012. [Online]. Available: https://doi.org/10.1145/2391229.2391236
2012
Earlier work this paper cites.
Intel, Hewlett-Packard, NEC, and Dell, “Intelligent platform management interface specification second generation v2.0,” Intel, Tech. Rep., 2013
2013
Earlier work this paper cites.
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman, “Project adam: Building an efficient and scalable deep learning training system,” in 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14) . Broomfield, CO: USENIX Association, Oct. 2014, pp. 571–582. [Online]. Available: https://www.usenix.org/conference/osdi14/technical-sessions/presentation/chilimbi
2014
Earlier work this paper cites.
D. Tiwari, S. Gupta, J. Rogers, D. Maxwell, P. Rech, S. Vazhkudai, D. Oliveira, D. Londo, N. DeBardeleben, P. Navaux, L. Carro, and A. Bland, “Understanding GPU errors on large-scale HPC systems and the implications for system design and operation,” in 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA) , 2015, pp. 331–342
2015
Earlier work this paper cites.
A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes, “Large-scale cluster management at google with borg,” in Proceedings of the Tenth European Conference on Computer Systems , ser. EuroSys ’15. New York, NY, USA: Association for Computing Machinery, 2015. [Online]. Available: https://doi.org/10.1145/2741948.2741964
2015
Earlier work this paper cites.
S. Levy, K. B. Ferreira, N. DeBardeleben, T. Siddiqua, V. Sridharan, and E. Baseman, “Lessons learned from memory errors observed over the lifetime of cielo,” in SC18: International Conference for High Performance Computing, Networking, Storage and Analysis , 2018, pp. 554–565
2018
Earlier work this paper cites.
C. Zimmer, D. Maxwell, S. McNally, S. Atchley, and S. S. Vazhkudai, “Gpu age-aware scheduling to improve the reliability of leadership jobs on titan,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis , ser. SC ’18. IEEE Press, 2019. [Online]. Available: https://doi.org/10.1109/SC.2018.00010
2018
Earlier work this paper cites.
M. Jeon, S. Venkataraman, A. Phanishayee, J. Qian, W. Xiao, and F. Yang, “Analysis of Large-Scale Multi-Tenant GPU clusters for DNN training workloads,” in 2019 USENIX Annual Technical Conference (USENIX ATC 19) . Renton, WA: USENIX Association, Jul. 2019, pp. 947–960. [Online]. Available: https://www.usenix.org/conference/atc19/presentation/jeon
2019
Earlier work this paper cites.
B. Roziere, M.-A. Lachaux, L. Chanussot, and G. Lample, “Unsupervised translation of programming languages,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Earlier work this paper cites.
R. Bonderson, “Training in turmoil: Silent data corruption in systems at scale,” 2021, for submission to an invited talk at the International Test Conference, in a ”Silicon Lifecycle Management” workshop. Conference is October 10-15, with this presentation / talk sometime on the 15th
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Maeng, S. Bharuka, I. Gao, M. C. Jeffrey, V. Saraph, B.-Y. Su, C. Trippel, J. Yang, M. Rabbat, B. Lucia, and C.-J. Wu, “Cpr: Understanding and improving failure tolerant training for deep learning recommendation with partial recovery,” in Proceedings of Machine Learning and Systems , 2021
2021
Earlier work this paper cites.
P. Barham, A. Chowdhery, J. Dean, S. Ghemawat, S. Hand, D. Hurt, M. Isard, H. Lim, R. Pang, S. Roy, B. Saeta, P. Schuh, R. Sepassi, L. E. Shafey, C. A. Thekkath, and Y. Wu, “Pathways: Asynchronous distributed dataflow for ml,” in Proceedings of Machine Learning and Systems , 2022
2022
Cited alongside, same era.
A. Benoit, Y. Du, T. Herault, L. Marchal, G. Pallez, L. Perotin, Y. Robert, H. Sun, and F. Vivien, “Checkpointing à la young/daly: An overview,” in Proceedings of the 2022 Fourteenth International Conference on Contemporary Computing , ser. IC3-2022. New York, NY, USA: Association for Computing Machinery, 2022, p. 701–710. [Online]. Available: https://doi.org/10.1145/3549206.3549328
2022
Cited alongside, same era.
2022
Cited alongside, same era.
“Nvidia dgx a100,” https://images.nvidia.com/aem-dam/Solutions/Data-Center/nvidia-dgx-a100-datasheet.pdf , (Accessed on 8/7/2024)
2024
Closest in time.
“Nvidia gb200 nvl72,” https://www.nvidia.com/en-us/data-center/gb200-nvl72/ , (Accessed on 9/13/2024)
2024
Closest in time.
“Nvidia xid error messages,” https://docs.nvidia.com/deploy/pdf/XID_Errors.pdf , (Accessed on 7/19/2024)
2024
Closest in time.
“Prolog and epilog guide,” https://slurm.schedmd.com/prolog_epilog.html , (Accessed on 8/7/2024)
2024
Closest in time.
“(prototype) flight recorder for debugging stuck jobs,” https://pytorch.org/tutorials/prototype/flight_recorder_tutorial.html , (Accessed on 11/4/2024)
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Li, R. Arora, S. Samsi, T. Patel, W. Arcand, D. Bestor, C. Byun, R. B. Roy, B. Bergeron, J. Holodnak, M. Houle, M. Hubbell, M. Jones, J. Kepner, A. Klein, P. Michaleas, J. McDonald, L. Milechin, J. Mullen, A. Prout, B. Price, A. Reuther, A. Rosa, M. Weiss, C. Yee, D. Edelman, A. Vanterpool, A. Cheng, V. Gadepally, and D. Tiwari, “AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , 2022, pp. 1224–1237
2022
Cited alongside, same era.
2022
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel, “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023. [Online]. Available: http://jmlr.org/papers/v24/22-1144.html
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. He, M. Hutton, S. Chan, R. De Gruijl, R. Govindaraju, N. Patil, and Y. Li, “Understanding and mitigating hardware failures in deep learning training systems,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , ser. ISCA ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3579371.3589105
2023
Cited alongside, same era.
N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Young, X. Zhou, Z. Zhou, and D. A. Patterson, “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“The shield: Self-healing interconnect,” https://network.nvidia.com/related-docs/whitepapers/WP_Mellanox_SHIELD.pdf , (Accessed on 7/18/2024)
2024
Closest in time.
“Submit it!” https://github.com/facebookincubator/submitit , (Accessed on 9/15/2024)
2024
Closest in time.
J. Ansel, E. Yang, H. He, N. Gimelshein, A. Jain, M. Voznesensky, B. Bao, P. Bell, D. Berard, E. Burovski, G. Chauhan, A. Chourdia, W. Constable, A. Desmaison, Z. DeVito, E. Ellison, W. Feng, J. Gong, M. Gschwind, B. Hirsh, S. Huang, K. Kalambarkar, L. Kirsch, M. Lazos, M. Lezcano, Y. Liang, J. Liang, Y. Lu, C. K. Luk, B. Maher, Y. Pan, C. Puhrsch, M. Reso, M. Saroufim, M. Y. Siraichi, H. Suk, S. Zhang, M. Suo, P. Tillet, X. Zhao, E. Wang, K. Zhou, R. Zou, X. Wang, A. Mathews, W. Wen, G. Chanan, P. Wu, and S. Chintala, “Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024
2024
Closest in time.
L. Bautista-Gomez, A. Benoit, S. Di, T. Herault, Y. Robert, and H. Sun, “A survey on checkpointing strategies: Should we always checkpoint à la young/daly?” Future Generation Computer Systems , vol. 161, pp. 315–328, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167739X24003777
2024
Closest in time.
A. Choudhury, Y. Wang, T. Pelkonen, K. Srinivasan, A. Jain, S. Lin, D. David, S. Soleimanifard, M. Chen, A. Yadav, R. Tijoriwala, D. Samoylov, and C. Tang, “MAST: Global scheduling of ML training across Geo-Distributed datacenters at hyperscale,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Association, Jul. 2024, pp. 563–580. [Online]. Available: https://www.usenix.org/conference/osdi24/presentation/choudhury
2024
Closest in time.
A. Douillard, Q. Feng, A. A. Rusu, R. Chhaparia, Y. Donchev, A. Kuncoro, M. Ranzato, A. Szlam, and J. Shen, “Diloco: Distributed low-communication training of language models,” in 2nd Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ICML 2024) , 2024. [Online]. Available: https://openreview.net/forum?id=pICSfWkJIk
2024
Closest in time.
A. Erben and E. Erdil, “Hardware failures won’t limit ai scaling,” 2024, accessed: 2024-12-06. [Online]. Available: https://epoch.ai/blog/hardware-failures-wont-limit-ai-scaling
2024
Closest in time.
2024
Closest in time.
S. Hsia, A. Golden, B. Acun, N. Ardalani, Z. DeVito, G.-Y. Wei, D. Brooks, and C.-J. Wu, “Mad-max beyond single-node: Enabling large machine learning model acceleration on distributed systems,” in 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , 2024, pp. 818–833
2024
Closest in time.
Q. Hu, Z. Ye, Z. Wang, G. Wang, M. Zhang, Q. Chen, P. Sun, D. Lin, X. Wang, Y. Luo, Y. Wen, and T. Zhang, “Characterization of large language model development in the datacenter,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) . Santa Clara, CA: USENIX Association, Apr. 2024, pp. 709–729. [Online]. Available: https://www.usenix.org/conference/nsdi24/presentation/hu
2024
Closest in time.
Z. Jiang, H. Lin, Y. Zhong, Q. Huang, Y. Chen, Z. Zhang, Y. Peng, X. Li, C. Xie, S. Nong, Y. Jia, S. He, H. Chen, Z. Bai, Q. Hou, S. Yan, D. Zhou, Y. Sheng, Z. Jiang, H. Xu, H. Wei, Z. Zhang, P. Nie, L. Zou, S. Zhao, L. Xiang, Z. Liu, Z. Li, X. Jia, J. Ye, X. Jin, and X. Liu, “MegaScale: Scaling large language model training to more than 10,000 GPUs,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) . Santa Clara, CA: USENIX Association, Apr. 2024, pp. 745–760. [Online]. Available: https://www.usenix.org/conference/nsdi24/presentation/jiang-ziheng
2024
Closest in time.
D. Ma, F. Lin, A. Desmaison, J. Coburn, D. Moore, S. Sankar, and X. Jiao, “Dr. dna: Combating silent data corruptions in deep learning using distribution of neuron activations,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , ser. ASPLOS ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 239–252. [Online]. Available: https://doi.org/10.1145/3620666.3651349
2024
Closest in time.
Meta, “The official Meta Llama 3 GitHub site,” 2024. [Online]. Available: https://github.com/meta-llama/llama3
2024
Closest in time.
2024
Closest in time.
K. Qian, Y. Xi, J. Caoa, J. Gao, Y. Xu, Y. Guan, B. Fu, X. Shi, F. Zhu, R. Miao, C. Wang, P. Wang, P. Zhang, X. Zeng, Z. Yao, E. Zhai, and D. Cai, “Alibaba hpn: A data center network for large language model training,” in SIGCOMM , 2024
2024
Closest in time.