Fetching the paper…
Reading the bibliography…
In order to satisfy their ever increasing capacity and compute requirements, machine learning models are distributed across multiple nodes using numerous parallelism strategies.
A study of persistent threads style gpu programming for gpgpu workloads
Kshitij Gupta, Jeff A. Stuart, and John D. Owens · 2012
Earlier work this paper cites.
A performance study to guide rdma programming decisions
Patrick MacArthur and Robert D. Russell · 2012
Earlier work this paper cites.
Cuda m3: Designing efficient cuda managed memory-aware mpi by exploiting gdr and ipc
Khaled Hamidouche, Ammar Ahmad Awan, Akshay Venkatesh, and Dhabaleswar K. Panda · 2016
Earlier work this paper cites.
High-performance key-value store on openshmem
Huansong Fu, Manjunath Gorentla Venkata, Ahana Roy Choudhury, Neena Imam, and Weikuan Yu · 2017
Earlier work this paper cites.
Locality-aware cta clustering for modern gpus
Ang Li, Shuaiwen Leon Song, Weifeng Liu, Xu Liu, Akash Kumar, and Henk Corporaal · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Performance evaluation of mpi libraries on gpu-enabled openpower architectures: Early experiences
Kawthar Shafie Khorassani, Ching-Hsiang Chu, Hari Subramoni, and Dhabaleswar K. Panda · 2019
Earlier work this paper cites.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherniavskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kondratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao, Bill Jia, Liang Xiong, and Misha Smelyanskiy · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, H. T. Kung, and David Cox · 2019
Earlier work this paper cites.
Nv-group: link-efficient reduction for distributed deep learning on modern dense gpu systems
Ching-Hsiang Chu, Pouya Kousha, Ammar Ahmad Awan, Kawthar Shafie Khorassani, Hari Subramoni, and Dhabaleswar K. (D K) Panda · 2020
Earlier work this paper cites.
GPU Initiated OpenSHMEM: Correct and Efficient Intra-Kernel Networking for DGPUs
Khaled Hamidouche and Michael LeBeane · 2020
Earlier work this paper cites.
ASTRA-SIM: Enabling SW/HW Co-Design Exploration for Distributed DL Training Platforms
Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna · 2020
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Cited alongside, same era.
https://engineering.fb.com/2021/07/15/open-source/fsdp/
Fully Sharded Data Parallel: faster AI training with fewer GPUs · 2021
Cited alongside, same era.
Synthesizing optimal collective algorithms
Zixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi, Todd Mytkowicz, Jacob Nelson, and Olli Saarikivi · 2021
Cited alongside, same era.
Communication algorithm-architecture co-design for distributed deep learning
Jiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid, Ki Hwan Yum, and Eun Jung Kim · 2021
Cited alongside, same era.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale, 2022
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Later among the works it cites.
Themis: A network bandwidth-aware collective scheduling policy for distributed training of dl models
Saeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan, and Tushar Krishna · 2022
Later among the works it cites.
Machine learning model sizes and the parameter gap, 2022
Pablo Villalobos, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, Anson Ho, and Marius Hobbhahn · 2022
Later among the works it cites.
Overlap communication with dependent computation via decomposition in large deep learning models
Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, Sameer Kumar, Tongfei Guo, Yuanzhong Xu, and Zongwei Zhou · 2022
Later among the works it cites.
Empowering gnns with fine-grained communication-computation pipelining on multi-gpu platforms, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimizing gpu kernel launch overhead in deep learning inference on mobile gpus
Sumin Kim, Seunghwan Oh, and Youngmin Yi · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Tutel: Adaptive mixture-of-experts at scale, 2022
Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, Joe Chau, Peng Cheng, Fan Yang, Mao Yang, and Yongqiang Xiong · 2022
Cited alongside, same era.
Breaking the computation and communication abstraction barrier in distributed machine learning workloads
Abhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet, Saeed Maleki, Youshan Miao, Madanlal Musuvathi, Todd Mytkowicz, and Olli Saarikivi · 2022
Cited alongside, same era.
Software-hardware co-design for fast and scalable training of deep learning recommendation models
Dheevatsa Mudigere, Yuchen Hao, Jianyu Huang, Zhihao Jia, Andrew Tulloch, Srinivas Sridharan, Xing Liu, Mustafa Ozdal, Jade Nie, Jongsoo Park, Liang Luo, Jie (Amy) Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang, Rakesh Komuravelli, Ching-Hsiang Chu, Serhat Yilmaz, Huayu Li, Jiyuan Qian, Zhuobo Feng, Yinbin Ma, Junjie Yang, Ellie Wen, Hong Li, Lin Yang, Chonglin Sun, Whitney Zhao, Dimitry Melts, Krishna Dhulipala, KR Kishore, Tyler Graf, Assaf Eisenman, Kiran Kumar Matam, Adi Gangidi, Guoqiang Jerry Chen, Manoj Krishnan, Avinash Nayak, Krishnakumar Nair, Bharath Muthiah, Mahmoud khorashadi, Pallab Bhattacharya, Petr Lapukhov, Maxim Naumov, Ajit Mathews, Lin Qiao, Mikhail Smelyanskiy, Bill Jia, and Vijay Rao · 2022
Cited alongside, same era.
https://pytorch.org/blog/accelerating-triton/
Accelerating Triton Dequantization Kernels for GPTQ
Cited in the paper.
https://www.amd.com/system/files/documents/amd-cdna2-white-paper.pdf
AMD CDNA™ 2 ARCHITECTURE
Cited in the paper.
Yuke Wang, Boyuan Feng, Zheng Wang, Tong Geng, Kevin Barker, Ang Li, and Yufei Ding · 2022
Later among the works it cites.
Fleche: An efficient gpu embedding cache for personalized recommendations
Minhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang, Jian Gao, Kai Ren, and Jiwu Shu · 2022
Later among the works it cites.
Ark: Gpu-driven code execution for distributed deep learning
Changho Hwang, KyoungSoo Park, Ran Shu, Xinyuan Qu, Peng Cheng, and Yongqiang Xiong · 2023
Closest in time.
A framework for fine-grained synchronization of dependent gpu kernels, 2023
Abhinav Jangda, Saeed Maleki, Maryam Mehri Dehnavi, Madan Musuvathi, and Olli Saarikivi · 2023
Closest in time.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini · 2023
Closest in time.
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives, 2024
Suchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena, and Matthew D. Sinclair · 2024
Closest in time.