Fetching the paper…
Reading the bibliography…
The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model.
C. Clos, “A study of non-blocking switching networks,” The Bell System Technical Journal , vol. 32, no. 2, pp. 406–424, 1953
1953
Earlier work this paper cites.
Dally and Seitz, “Deadlock-free message routing in multiprocessor interconnection networks,” IEEE Transactions on Computers , vol. C-36, no. 5, pp. 547–553, 1987
1987
Earlier work this paper cites.
R. Koo and S. Toueg, “Checkpointing and rollback-recovery for distributed systems,” IEEE Transactions on software Engineering , no. 1, pp. 23–31, 1987
1987
Earlier work this paper cites.
L. G. Valiant, “A bridging model for parallel computation,” Commun. ACM , vol. 33, no. 8, p. 103–111, aug 1990. [Online]. Available: https://doi.org/10.1145/79173.79181
1990
Earlier work this paper cites.
C. Hopps, “Rfc2992: Analysis of an equal-cost multi-path algorithm,” 2000
2000
Earlier work this paper cites.
S. S. Mukherjee, M. Kontz, and S. K. Reinhardt, “Detailed design and evaluation of redundant multithreading alternatives,” ACM SIGARCH Computer Architecture News , vol. 30, no. 2, pp. 99–110, 2002
2002
Earlier work this paper cites.
H. Weatherspoon and J. D. Kubiatowicz, “Erasure coding vs. replication: A quantitative comparison,” in International Workshop on Peer-to-Peer Systems . Springer, 2002, pp. 328–337
2002
Earlier work this paper cites.
M. K. Aguilera, R. Janakiraman, and L. Xu, “Using erasure codes efficiently for storage in a distributed system,” in 2005 International Conference on Dependable Systems and Networks (DSN’05) . IEEE, 2005, pp. 336–345
2005
Earlier work this paper cites.
A. G. Dimakis, V. Prabhakaran, and K. Ramchandran, “Decentralized erasure codes for distributed networked storage,” IEEE Transactions on Information Theory , vol. 52, no. 6, pp. 2809–2816, 2006
2006
Earlier work this paper cites.
H. Han, S. Shakkottai, C. V. Hollot, R. Srikant, and D. Towsley, “Multi-path tcp: A joint congestion control and routing scheme to exploit path diversity in the internet,” IEEE/ACM Transactions on Networking , vol. 14, no. 6, pp. 1260–1271, 2006
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
J. C. Smolens, B. T. Gold, B. Falsafi, and J. C. Hoe, “Reunion: Complexity-effective multicore redundancy,” in 2006 39th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO’06) , 2006, pp. 223–234
2006
Earlier work this paper cites.
S. Poledna, Fault-tolerant real-time systems: The problem of replica determinism . Springer Science & Business Media, 2007, vol. 345
2007
Earlier work this paper cites.
M. Kutare, G. Eisenhauer, C. Wang, K. Schwan, V. Talwar, and M. Wolf, “Monalytics: online monitoring and analytics for managing large scale data centers,” in Proceedings of the 7th international conference on Autonomic computing , 2010, pp. 141–150
2010
Earlier work this paper cites.
A. Dixit, P. Prakash, Y. C. Hu, and R. R. Kompella, “On the impact of packet spraying in data center networks,” in 2013 Proceedings IEEE INFOCOM , 2013, pp. 2130–2138
2013
Earlier work this paper cites.
M. Xia, M. Saxena, M. Blaum, and D. A. Pease, “A tale of two erasure codes in HDFS,” in 13th USENIX Conference on File and Storage Technologies (FAST 15) . Santa Clara, CA: USENIX Association, Feb. 2015, pp. 213–226. [Online]. Available: https://www.usenix.org/conference/fast15/technical-sessions/presentation/xia
2015
Earlier work this paper cites.
C. Yu, C. Lumezanu, A. Sharma, Q. Xu, G. Jiang, and H. V. Madhyastha, “Software-defined latency monitoring in data center networks,” in Passive and Active Measurement: 16th International Conference, PAM 2015, New York, NY, USA, March 19-20, 2015, Proceedings 16 . Springer, 2015, pp. 360–372
2015
Earlier work this paper cites.
S. Jeaugey, “Nccl 2.0,” in GPU Technology Conference (GTC) , vol. 2, 2017
2017
Cited alongside, same era.
R. Joshi, T. Qu, M. C. Chan, B. Leong, and B. T. Loo, “Burstradar: Practical real-time microburst monitoring for datacenter networks,” in Proceedings of the 9th Asia-Pacific Workshop on Systems , 2018, pp. 1–8
2018
Cited alongside, same era.
H. Li, “Alluxio: A virtual distributed file system,” Ph.D. dissertation, EECS Department, University of California, Berkeley, May 2018. [Online]. Available: http://www2.eecs.berkeley.edu/Pubs/TechRpts/2018/EECS-2018-29.html
2018
Cited alongside, same era.
Y. Lu, G. Chen, B. Li, K. Tan, Y. Xiong, P. Cheng, J. Zhang, E. Chen, and T. Moscibroda, “Multi-Path transport for RDMA in datacenters,” in 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18) . Renton, WA: USENIX Association, Apr. 2018, pp. 357–371. [Online]. Available: https://www.usenix.org/conference/nsdi18/presentation/lu
2018
Cited alongside, same era.
2021
Later among the works it cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” 2022
2022
Later among the works it cites.
A. Eisenman, K. K. Matam, S. Ingram, D. Mudigere, R. Krishnamoorthi, K. Nair, M. Smelyanskiy, and M. Annavaram, “ { \{ Check-N-Run } \} : A checkpointing system for training deep learning recommendation models,” in 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) , 2022, pp. 929–943
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Cited alongside, same era.
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
J. Dong, Z. Cao, T. Zhang, J. Ye, S. Wang, F. Feng, L. Zhao, X. Liu, L. Song, L. Peng, Y. Guo, X. Jiang, L. Tang, Y. Du, Y. Zhang, P. Pan, and Y. Xie, “Eflops: Algorithm and system co-design for a high performance distributed training platform,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020, pp. 610–622
2020
Cited alongside, same era.
B. Nicolae, J. Li, J. M. Wozniak, G. Bosilca, M. Dorier, and F. Cappello, “Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models,” in 2020 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID) . IEEE, 2020, pp. 172–181
2020
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” 2020
2020
Cited alongside, same era.
2023
Later among the works it cites.
M. Cowan, S. Maleki, M. Musuvathi, O. Saarikivi, and Y. Xiong, “Mscclang: Microsoft collective communication language,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 502–514
2023
Later among the works it cites.
W. Cui, Z. Han, L. Ouyang, Y. Wang, N. Zheng, L. Ma, Y. Yang, F. Yang, J. Xue, L. Qiu et al. , “Optimizing dynamic neural networks with brainstorm,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) , 2023, pp. 797–815
2023
Later among the works it cites.
Facebook, “Gloo,” 2023, https://github.com/facebookincubator/gloo
2023
Later among the works it cites.
J. Liu, J. H. Wang, and Y. Jiang, “Janus: A unified distributed training framework for sparse mixture-of-experts models,” in Proceedings of the ACM SIGCOMM 2023 Conference , 2023, pp. 486–498
2023
Later among the works it cites.
Microsoft, “Nsccl,” 2023, https://github.com/microsoft/msccl
2023
Later among the works it cites.
Nvidia, “Nvidia nsight systems,” 2023, https://developer.nvidia.com/nsight-systems
2023
Later among the works it cites.
A. Shah, V. Chidambaram, M. Cowan, S. Maleki, M. Musuvathi, T. Mytkowicz, J. Nelson, O. Saarikivi, and R. Singh, “ { \{ TACCL } \} : Guiding collective algorithm synthesis using communication sketches,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , 2023, pp. 593–612
2023
Later among the works it cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Later among the works it cites.
Z. Wang, Z. Jia, S. Zheng, Z. Zhang, X. Fu, T. E. Ng, and Y. Wang, “Gemini: Fast failure recovery in distributed training with in-memory checkpoints,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 364–381
2023
Later among the works it cites.
——, “Gemini: Fast failure recovery in distributed training with in-memory checkpoints,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 364–381
2023
Later among the works it cites.
C. Zhang, L. Ma, J. Xue, Y. Shi, Z. Miao, F. Yang, J. Zhai, Z. Yang, and M. Yang, “Cocktailer: Analyzing and optimizing dynamic control flow in deep learning,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) , 2023, pp. 681–699
2023
Later among the works it cites.
2024
Closest in time.