Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) large language models (LLM) have memory requirements that often exceed the GPU memory capacity, requiring costly parameter movement from secondary memories to the GPU for expert computation.
Ramulator: A fast and extensible DRAM simulator
Yoongu Kim et al · 2015
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In ICLR
Noam Shazeer et al · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT
Jacob Devlin et al · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel et al · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD
Jeff Rasley et al · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing. In EMNLP
Thomas Wolf et al · 2020
Earlier work this paper cites.
Hugging Face Switch Transformers
Google. 2021 · 2021
Cited alongside, same era.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. In SC
Samyam Rajbhandari et al · 2021
Cited alongside, same era.
ZeRO-Offload: Democratizing Billion-Scale Model Training.. In ATC
Jie Ren et al · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus et al · 2022
Cited alongside, same era.
Hugging Face NLLB MoE Model Hub
Meta. 2022 · 2022
Cited alongside, same era.
Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference
Haiyang Huang et al · 2023
Later among the works it cites.
Samsung PIM/PNM for Transformer Based AI: Energy Efficiency on PIM/PNM Cluster. In HCS
Jin Hyun Kim et al · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Swapnil Sharma et al · 2023
Later among the works it cites.
SE-MoE: A Scalable and Efficient Mixture-of-Experts Distributed Training and Inference System
Liang Shen et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…