Fetching the paper…
Reading the bibliography…
Neural networks have become a cornerstone of machine learning.
“Placeto: Learning Generalizable Device Placement Algorithms for Distributed Machine Learning”, 2019
Ravichandra Addanki et al · 1906
Earlier work this paper cites.
“Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”, 2020
Mohammad Shoeybi et al · 1909
Earlier work this paper cites.
“Parallel static and dynamic multi-constraint graph partitioning”
Kirk Schloegel, George Karypis and Vipin Kumar · 2002
Earlier work this paper cites.
“Training Deep Nets with Sublinear Memory Cost”, 2016
Tianqi Chen, Bing Xu, Chiyuan Zhang and Carlos Guestrin · 2016
Earlier work this paper cites.
“Graph Partitioning with Acyclicity Constraints”, 2017
Orlando Moreira, Merten Popp and Christian Schulz · 2017
Earlier work this paper cites.
“PipeDream: Fast and Efficient Pipeline Parallel DNN Training”, 2018
Aaron Harlap et al · 2018
Earlier work this paper cites.
“Exploring Hidden Dimensions in Accelerating Convolutional Neural Networks”
Zhihao Jia, Sina Lin, Charles. Qi and Alex Aiken · 2018
Earlier work this paper cites.
“Evolutionary multi-level acyclic graph partitioning”
Orlando Moreira, Merten Popp and Christian Schulz · 2018
Earlier work this paper cites.
“Mesh-TensorFlow: Deep Learning for Supercomputers”
Noam Shazeer et al · 2018
Earlier work this paper cites.
“GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism”
Yanping Huang et al · 2019
Earlier work this paper cites.
“Beyond Data and Model Parallelism for Deep Neural Networks.”
Zhihao Jia, Matei Zaharia and Alex Aiken · 2019
Earlier work this paper cites.
“PipeDream: generalized pipeline parallelism for DNN training”
Deepak Narayanan et al · 2019
Earlier work this paper cites.
“Supporting Very Large Models using Automatic Dataflow Graph Partitioning”
Minjie Wang, Chien-chin Huang and Jinyang Li · 2019
Cited alongside, same era.
“Language Models are Few-Shot Learners”
Tom Brown et al · 2020
Cited alongside, same era.
“DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters”
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase and Yuxiong He · 2020
Cited alongside, same era.
“TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism”, 2022
Zhenkun Cai et al · 2021
Cited alongside, same era.
“NVIDIA A100 Tensor Core GPU: Performance and Innovation”, 2021
Jack Choquette et al · 2021
Cited alongside, same era.
“Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism”
“NAS-Bench-Suite: NAS Evaluation is (Now) Surprisingly Easy”, 2022
Yash Mehta et al · 2022
Later among the works it cites.
“Scaling Language Models: Methods, Analysis & Insights from Training Gopher”, 2022
Jack. Rae et al · 2022
Later among the works it cites.
“Compute Trends Across Three Eras of Machine Learning”
Jaime Sevilla et al · 2022
Later among the works it cites.
“Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model”, 2022
Shaden Smith et al · 2022
Later among the works it cites.
“Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization”
Colin Unger et al · 2022
Later among the works it cites.
“Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saar Eliad et al · 2021
Cited alongside, same era.
“HW-NAS-Bench:Hardware-Aware Neural Architecture Search Benchmark”, 2021
Chaojian Li et al · 2021
Cited alongside, same era.
“Efficient large-scale language model training on GPU clusters using megatron-LM”
Deepak Narayanan et al · 2021
Cited alongside, same era.
“Automatic Graph Partitioning for Very Large-scale Deep Learning”
Masahiro Tanaka, Kenjiro Taura, Toshihiro Hanawa and Kentaro Torisawa · 2021
Cited alongside, same era.
“Efficient and Systematic Partitioning of Large and Deep Neural Networks for Parallelization”
Haoran Wang et al · 2021
Cited alongside, same era.
“GSPMD: General and Scalable Parallelization for ML Computation Graphs”, 2021
Yuanzhong Xu et al · 2021
Cited alongside, same era.
“Pathways: Asynchronous Distributed Dataflow for ML”
Paul Barham et al · 2022
Cited alongside, same era.
Lianmin Zheng et al · 2022
Later among the works it cites.
“PaLM 2 Technical Report”, 2023
Rohan Anil et al · 2023
Later among the works it cites.
“PaLM: Scaling Language Modeling with Pathways”, 2023
Aakanksha Chowdhery et al · 2023
Later among the works it cites.
“Reducing activation recomputation in large transformer models”
Vijay Korthikanti et al · 2023
Later among the works it cites.
“A Survey on Auto-Parallelism of Large-Scale Deep Learning Training”, 2023
Peng Liang et al · 2023
Later among the works it cites.
“GPT-4 Technical Report”, 2023
OpenAI et al · 2023
Later among the works it cites.