Fetching the paper…
Reading the bibliography…
Large-scale distributed training in production datacenters constitutes a challenging workload bottlenecked by network communication.
Hedera: Dynamic flow scheduling for data center networks
Mohammad Al-Fares, Sivasankar Radhakrishnan, Barath Raghavan, Nelson Huang, and Amin Vahdat · 2010
Earlier work this paper cites.
Data center tcp (dctcp)
Mohammad Alizadeh, Albert Greenberg, David A. Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan · 2010
Earlier work this paper cites.
Multipath tcp: From theory to practice
Sébastien Barré, Christoph Paasch, and Olivier Bonaventure · 2011
Earlier work this paper cites.
Design, implementation and evaluation of congestion control for multipath TCP
Damon Wischik, Costin Raiciu, Adam Greenhalgh, and Mark Handley · 2011
Earlier work this paper cites.
Conga: Distributed congestion-aware load balancing for datacenters
Mohammad Alizadeh, Tom Edsall, Sarang Dharmapurikar, Ramanan Vaidyanathan, Kevin Chu, Andy Fingerhut, Vinh The Lam, Francis Matus, Rong Pan, Navindra Yadav, and George Varghese · 2014
Earlier work this paper cites.
Multipath tcp
Christoph Paasch and Olivier Bonaventure · 2014
Earlier work this paper cites.
Timely: Rtt-based congestion control for the datacenter
Radhika Mittal, Vinh The Lam, Nandita Dukkipati, Emily Blem, Hassan Wassel, Monia Ghobadi, Amin Vahdat, Yaogong Wang, David Wetherall, and David Zats · 2015
Earlier work this paper cites.
Congestion control for large-scale rdma deployments
Yibo Zhu, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, and Ming Zhang · 2015
Earlier work this paper cites.
Traffic engineering with equal-cost-multipath: An algorithmic perspective
Marco Chiesa, Guy Kindler, and Michael Schapira · 2016
Earlier work this paper cites.
Hula: Scalable load balancing using programmable data planes
Naga Katta, Mukesh Hira, Changhoon Kim, Anirudh Sivaraman, and Jennifer Rexford · 2016
Earlier work this paper cites.
Drill: Micro load balancing for low-latency data center networks
Soudeh Ghorbani, Zibin Yang, P. Brighten Godfrey, Yashar Ganjali, and Amin Firoozshahian · 2017
Earlier work this paper cites.
Re-architecting datacenter networks and stacks for low latency and high performance
Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew W. Moore, Gianni Antichi, and Marcin Wójcik · 2017
Earlier work this paper cites.
Hpcc: High precision congestion control
Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang, Zheng Cao, Ming Zhang, Frank Kelly, Mohammad Alizadeh, and Minlan Yu · 2019
Earlier work this paper cites.
Pipedream: generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, Phillip B. Gibbons, and Matei Zaharia · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Swift: Delay is simple and effective for congestion control in the datacenter
Gautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel, Xian Wu, Behnam Montazeri, Yaogong Wang, Kevin Springborn, Christopher Alfeld, Michael Ryan, David Wetherall, and Amin Vahdat · 2020
Cited alongside, same era.
Astra-sim: Enabling sw/hw co-design exploration for distributed dl training platforms
Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna · 2020
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using megatron-lm
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia · 2021
TACCL: Guiding collective algorithm synthesis using communication sketches
Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musuvathi, Todd Mytkowicz, Jacob Nelson, Olli Saarikivi, and Rachee Singh · 2023
Later among the works it cites.
TopoOpt: Co-optimizing network topology and parallelization strategy for distributed training jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Kewitsch · 2023
Later among the works it cites.
Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale
William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna · 2023
Later among the works it cites.
Smartt-reps: Sender-based marked rapidly-adapting trimmed & timed transport with recycled entropies
Tommaso Bonato, Abdul Kabbani, Daniele De Sensi, Rong Pan, Yanfang Le, Costin Raiciu, Mark Handley, Timo Schneider, Nils Blach, Ahmad Ghalayini, et al · 2024
Closest in time.
Reps: Recycling entropies for packet spraying to adaptively explore paths and mitigate failures
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hashing linearity enables relative path control in data centers
Zhehui Zhang, Haiyang Zheng, Jiayao Hu, Xiangning Yu, Chenchen Qi, Xuemei Shi, and Guohui Wang · 2021
Cited alongside, same era.
PowerTCP: Pushing the performance limits of datacenter networks
Vamsi Addanki, Oliver Michel, and Stefan Schmid · 2022
Cited alongside, same era.
Dcpim: Near-optimal proactive datacenter transport
Qizhe Cai, Mina Tahmasbi Arashloo, and Rachit Agarwal · 2022
Cited alongside, same era.
Backpressure flow control
Prateesh Goyal, Preey Shah, Kevin Zhao, Georgios Nikolaidis, Mohammad Alizadeh, and Thomas E. Anderson · 2022
Cited alongside, same era.
An edge-queued datagram service for all datacenter traffic
Vladimir Andrei Olteanu, Haggai Eran, Dragos Dumitrescu, Adrian Popa, Cristi Baciu, Mark Silberstein, Georgios Nikolaidis, Mark Handley, and Costin Raiciu · 2022
Cited alongside, same era.
An edge-queued datagram service for all datacenter traffic
Vladimir Andrei Olteanu, Haggai Eran, Dragos Dumitrescu, Adrian Popa, Cristi Baciu, Mark Silberstein, Georgios Nikolaidis, Mark Handley, and Costin Raiciu · 2022
Cited alongside, same era.
Efficient direct-connect topologies for collective communications
Liangyu Zhao, Siddharth Pal, Tapan Chugh, Weiyang Wang, Jason Fantl, Prithwish Basu, Joud Khoury, and Arvind Krishnamurthy · 2022
Cited alongside, same era.
Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Mohammad Dohadwala, Michael Papamichael, Daniele De Sensi, and Torsten Hoefler · 2024
Closest in time.
When ml training cuts through congestion: Just-in-time gradient compression via packet trimming
Xiaoqi Chen, Shay Vargaftik, and Ran Ben Basat · 2024
Closest in time.
Rdma over ethernet for distributed training at meta scale
Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, Shuqiang Zhang, Mikel Jimenez Fernandez, Shashidhar Gandham, and Hongyi Zeng · 2024
Closest in time.
MegaScale: Scaling large language model training to more than 10,000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi Zou, Sida Zhao, Liang Xiang, Zherui Liu, Zhe Li, Xiaoying Jia, Jianxi Ye, Xin Jin, and Xin Liu · 2024
Closest in time.
Alibaba hpn: A data center network for large language model training
Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, Chao Wang, Peng Wang, Pengcheng Zhang, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, and Dennis Cai · 2024
Closest in time.
CASSINI: Network-Aware job scheduling in machine learning clusters
Sudarsanan Rajasekaran, Manya Ghobadi, and Aditya Akella · 2024
Closest in time.
Mltcp: A distributed technique to approximate centralized flow scheduling for machine learning
Sudarsanan Rajasekaran, Sanjoli Narang, Anton A. Zabreyko, and Manya Ghobadi · 2024
Closest in time.
Towards Domain-Specific network transport for distributed DNN training
Hao Wang, Han Tian, Jingrong Chen, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai, Junchen Jiang, Yong Wang, and Kai Chen · 2024
Closest in time.