Fetching the paper…
Reading the bibliography…
Foundation models (e.g., ChatGPT, DALL-E, PengCheng Mind, PanGu-$\Sigma$) have demonstrated extraordinary performance in key technological areas, such as natural language processing and visual recognition, and have become the mainstream trend of artificial general intelligence.
Jacobs R A, et al. ”Adaptive mixtures of local experts.” Neural computation, 1991
1991
Earlier work this paper cites.
Van De Geijn R A, et al. ”SUMMA: Scalable universal matrix multiplication algorithm.” Concurrency: Practice and Experience, 1997
1997
Earlier work this paper cites.
Solomonik E, et al. ”Communication-optimal parallel 2.5 D matrix multiplication and LU factorization algorithms.” European Conference on Parallel Processing. Berlin, Heidelberg: Springer Berlin Heidelberg, 2011
2011
Earlier work this paper cites.
Chen T, et al. ”Training deep nets with sublinear memory cost.” arXiv, 2016
2016
Earlier work this paper cites.
Rhu M, et al. ”vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design.”,MICRO, 2016
2016
Earlier work this paper cites.
Vaswani A, et al. ”Attention is all you need.” NeurIP, 2017
2017
Earlier work this paper cites.
Gibiansky, et al. ”Bringing HPC techniques to deep learning.” Baidu Research, Tech. Rep. (2017)
2017
Earlier work this paper cites.
Devlin, et al. ”Bert: Pre-training of deep bidirectional transformers for language understanding.” NAACL,2018
2018
Earlier work this paper cites.
Micikevicius P, et al. ”Mixed precision training.”, ICLR, 2018
2018
Earlier work this paper cites.
Jia X, et al. ”Highly scalable deep learning training system with mixed-precision: Training imagenet in four minutes.” arXiv, 2018
2018
Earlier work this paper cites.
You Y, et al. ”Imagenet training in minutes.”, ICPP, 2018
2018
Earlier work this paper cites.
Huang Y, et al. ”Gpipe: Efficient training of giant neural networks using pipeline parallelism.”, NeurIPS, 2019
2019
Earlier work this paper cites.
Narayanan D, et al. ”PipeDream: generalized pipeline parallelism for DNN training.”, SOSP, 2019
2019
Earlier work this paper cites.
Paszke A, et al. ”Pytorch: An imperative style, high-performance deep learning library.”, NeurIPS, 2019
2019
Earlier work this paper cites.
Brown T, et al. ”Language models are few-shot learners.”, NeurIPS, 2020
2020
Earlier work this paper cites.
Li S, et al. ”PyTorch distributed: experiences on accelerating data parallel training.” VLDB, 2020
2020
Earlier work this paper cites.
Xu Y, et al. ”Automatic cross-replica sharding of weight update in data-parallel training.” arXiv, 2020
2020
Earlier work this paper cites.
Rasley J, et al. ”Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.”, KDD, 2020
2020
Earlier work this paper cites.
Huang C C, et al. ”Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping.”, ASPLOS, 2020
2020
Earlier work this paper cites.
Hildebrand M, et al. ”Autotm: Automatic tensor movement in heterogeneous memory systems using integer linear programming.”, ASPLOS, 2020
2020
Earlier work this paper cites.
Rajbhandari S, et al. ”Zero: Memory optimizations toward training trillion parameter models.” SC, 2020
2020
Earlier work this paper cites.
Huang C C, et al. ”Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping.”, ASPLOS, 2020
2020
Earlier work this paper cites.
Gujarati A, et al. ”Serving DNNs like Clockwork: Performance Predictability from the Bottom Up”, OSDI, 2020
2020
Earlier work this paper cites.
Crankshaw D, et al. ”InferLine: latency-aware provisioning and scaling for prediction serving pipelines.”, SoCC, 2020
2020
Earlier work this paper cites.
Wang R, et al. ”K-adapter: Infusing knowledge into pre-trained models with adapters.” arXiv, 2020
2020
Earlier work this paper cites.
Narayanan D, et al. ”Efficient large-scale language model training on gpu clusters using megatron-lm.”, SC, 2021
2021
Earlier work this paper cites.
Bian Z, et al. ”Maximizing parallelism in distributed training for huge neural networks.” arXiv, 2021
2021
Earlier work this paper cites.
Fan S, et al. ”DAPPLE: A pipelined data parallel approach for training large models.”, PPoPP, 2021
2021
Earlier work this paper cites.
Narayanan D, et al. ”Memory-efficient pipeline-parallel dnn training.”, ICML, 2021
2021
Earlier work this paper cites.
Li S, et al. ”Chimera: efficiently training large-scale neural networks with bidirectional pipelines.”, SC, 2021
2021
Earlier work this paper cites.
Eliad S, et al. ”Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism.”, ATC 2021
2021
Earlier work this paper cites.
Lepikhin D, et al. ”Gshard: Scaling giant models with conditional computation and automatic sharding.”, ICLR, 2021
2021
Earlier work this paper cites.
He J, et al. ”Fastmoe: A fast mixture-of-expert training system.”, arXiv, 2021
2021
Earlier work this paper cites.
Bae J, et al. ”FlashNeuron:SSD-Enabled Large-Batch Training of Very Deep Neural Networks.”, FAST, 2021
2021
Cited alongside, same era.
Ren J, et al. ”ZeRO-Offload: Democratizing Billion-Scale model training.”, ATC, 2021
2021
Cited alongside, same era.
Rajbhandari S, et al. ”Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning.”,SC, 2021
2021
Cited alongside, same era.
Gan S, et al. ”Bagua: scaling up distributed learning with system relaxations.”, arXiv,2021
2021
Cited alongside, same era.
Romero F, et al. ”INFaaS: Automated Model-less Inference Serving.”, ATC, 2021
2021
Cited alongside, same era.
Lin Z, et al. ”The adapter-bot: All-in-one controllable conversational model.”, AAAI, 2021
2021
Xu Q, et al. ”An efficient 2d method for training super-large deep learning models.”, IPDPS, 2023
2023
Later among the works it cites.
Liu Z, et al. ”Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency.”, SC, 2023
2023
Later among the works it cites.
Zhang W, et al. ”MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale Models.” DAC, 2023
2023
Later among the works it cites.
Chen Z, et al. ”Elastic Averaging for Efficient Pipelined DNN Training.”, PPoPP, 2023
2023
Later among the works it cites.
Kim T, et al. ”BPIPE: memory-balanced pipeline parallelism for training large language models.”, ICML, 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wang B, et al. ”Tesseract: Parallelize the Tensor Parallelism Efficiently.”, ICPP, 2022
2022
Cited alongside, same era.
Athlur S, et al. ”Varuna: scalable, low-cost training of massive deep learning models.”, EuroSys, 2022
2022
Cited alongside, same era.
He J, et al. ”FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models.”, PPoPP, 2022
2022
Cited alongside, same era.
Smith S, et al. ”Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model.” arXiv, 2022
2022
Cited alongside, same era.
Zheng L, et al. ”Alpa: Automating inter-and Intra-Operator parallelism for distributed deep learning.”, OSDI, 2022
2022
Cited alongside, same era.
Miao X, et al. ”Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism.”, VLDB, 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
Zhai M, et al. ”SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization.”, ATC 2023
2023
Later among the works it cites.
Li J, et al. ”Accelerating Distributed MoE Training and Inference with Lina.”, ATC, 2023
2023
Later among the works it cites.
Liu J, et al. ”Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts Models.”, SIGCOMM, 2023
2023
Later among the works it cites.
Jung J, et al. ”DeepUM: Tensor Migration and Prefetching in Unified Memory.”, ASPLOS, 2023
2023
Later among the works it cites.
Zhang H, et al. ”G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations.”, MICRO, 2023
2023
Later among the works it cites.
Wang S, et al. ”Overlap communication with dependent computation via decomposition in large deep learning models.”, ASPLOS, 2023
2023
Later among the works it cites.
Feng Y, et al. ”Mobius: Fine tuning large-scale models on commodity gpu servers.”,ASPLOS, 2023
2023
Later among the works it cites.
Wang G, et al. ”ZeRO++: Extremely Efficient Collective Communication for Giant Model Training.” arXiv, 2023
2023
Later among the works it cites.
Wang J, et al. ”CocktailSGD: Fine-tuning foundation models over 500Mbps networks.”, ICML, 2023
2023
Later among the works it cites.
Song J, et al. ”Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication Compression.”,ASPLOS, 2023
2023
Later among the works it cites.
Dai S, et al. ”Efficient Transformer Inference with Statically Structured Sparse Attention.”, DAC, 2023
2023
Later among the works it cites.
Liu Z, et al. ”Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time.”, ICML, 2023
2023
Later among the works it cites.
Zhang Z, et al. ”H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models”, ICML, 2023
2023
Later among the works it cites.
Guo L, et al. ”STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining.”,ASPLOS, 2023
2023
Later among the works it cites.
Guo C, et al. ”OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization.”, ISCA, 2023
2023
Later among the works it cites.
Li Z, et al. ”AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving.”, OSDI, 2023
2023
Later among the works it cites.
Wu B, et al. ”Fast Distributed Inference Serving for Large Language Models.” arXiv, 2023
2023
Later among the works it cites.
Zhang H, et al. ”SHEPHERD: Serving DNNs in the Wild.”, NSDI, 2023
2023
Later among the works it cites.
Liu L, et al. ”Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning Approach.”, FAST, 2023
2023
Later among the works it cites.
Sheng Y, et al. ”FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.”, ICML, 2023
2023
Later among the works it cites.
Jeong J, et al. ”Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access.”, EuroSys, 2023
2023
Later among the works it cites.
Kwon W, et al. ”Efficient memory management for large language model serving with pagedattention.”,SOSP, 2023
2023
Later among the works it cites.
Wang Y, et al. ”Tabi: An Efficient Multi-Level Inference System for Large Language Models.”, EuroSys, 2023
2023
Later among the works it cites.
Leviathan Y, et al. ”Fast Inference from Transformers via Speculative Decoding”, ICML, 2023
2023
Later among the works it cites.
Xu D, et al. ”LLMCad: Fast and Scalable On-device Large Language Model Inference.”, arxiv, 2023
2023
Later among the works it cites.
Jiang C, et al. ”DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines.”, EuroSys, 2024
2024
Closest in time.