Fetching the paper…
Reading the bibliography…
The Mixture-of-Experts (MoE) model has emerged as a prominent architecture in the field of Large Language Models (LLMs), providing a better balance between model performance and computational efficiency.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding, 2020
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Earlier work this paper cites.
Distributed hierarchical gpu parameter server for massive scale deep learning ads systems, 2020
Weijie Zhao, Deping Xie, Ronglai Jia, Yulei Qian, Ruiquan Ding, Mingming Sun, and Ping Li · 2020
Earlier work this paper cites.
Fastmoe: A fast mixture-of-expert training system, 2021
Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, and Jie Tang · 2021
Earlier work this paper cites.
M6-t: Exploring sparse expert models and beyond, 2021
An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, Di Zhang, Wei Lin, Lin Qu, Jingren Zhou, and Hongxia Yang · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Earlier work this paper cites.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production, 2022
Young Jin Kim, Rawn Henry, Raffy Fahim, and Hany Hassan Awadalla · 2022
Earlier work this paper cites.
Hetumoe: An efficient trillion-scale mixture-of-expert distributed training system, 2022
Xiaonan Nie, Pinxue Zhao, Xupeng Miao, Tong Zhao, and Bin Cui · 2022
Earlier work this paper cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale, 2022
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Earlier work this paper cites.
St-moe: Designing stable and transferable sparse expert models, 2022
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Earlier work this paper cites.
Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement
Xiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang, Jilong Xue, Lingxiao Ma, Gang Cao, and Bin Cui · 2023
Earlier work this paper cites.
A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training
Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan, Samyam Rajbhandari, Yuxiong He, and Abhinav Bhatele · 2023
Earlier work this paper cites.
https://docs.nvidia.com/cuda/cublas/
Cublas · 2024
Cited alongside, same era.
https://developer.nvidia.com/blog/cuda-pro-tip-increase-performance-with-vectorized-memory-access/
Cuda pro tip: Increase performance with vectorized memory access · 2024
Cited alongside, same era.
https://github.com/NVIDIA/cutlass/tree/main
Cutlass · 2024
Cited alongside, same era.
https://x.ai/blog/grok-2
Grok · 2024
Cited alongside, same era.
https://huggingface.co/mistralai/Mixtral-8x7B-v0.1
Mixtral · 2024
Cited alongside, same era.
https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/features/parallelisms.html
Nvidia parallelisms · 2024
Cited alongside, same era.
Flux: Fast software-based communication overlap on gpus through kernel fusion, 2024
Li-Wen Chang, Wenlei Bao, Qi Hou, Chengquan Jiang, Ningxin Zheng, Yinmin Zhong, Xuanrun Zhang, Zuquan Song, Ziheng Jiang, Haibin Lin, Xin Jin, and Xin Liu · 2024
Closest in time.
Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan · 2024
Closest in time.
P/d-serve: Serving disaggregated large language model at scale, 2024
Yibo Jin, Tao Wang, Huimin Lin, Mingyang Song, Peiyang Li, Yipeng Ma, Yicheng Shan, Zhengfan Yuan, Cailong Li, Yajing Sun, Tiandeng Wu, Xing Chu, Ruizhi Huan, Li Ma, Xiao You, Wenting Zhou, Yunpeng Ye, Wen Liu, Xiangkun Xu, Yongsheng Zhang, Tiantian Dong, Jiawei Zhu, Zhe Wang, Xijian Ju, Jianxun Song, Haoliang Cheng, Xiaojing Li, Jiandong Ding, Hefei Guo, and Zhengyong Zhang · 2024
Closest in time.
Scaling laws for fine-grained mixture of experts, 2024
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, Krystian Król, Tomasz Odrzygóźdź, Piotr Sankowski, Marek Cygan, and Sebastian Jaszczur · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
https://www.snowflake.com/en/blog/arctic-open-efficient-foundation-language-models-snowflake/
Snowflake arctic: The best llm for enterprise ai — efficiently intelligent, truly open · 2024
Cited alongside, same era.
https://huggingface.co/xverse/XVERSE-MoE-A36B
Xverse-moe-a36b · 2024
Cited alongside, same era.
Technical report, Google DeepMind, 2024
Gemini1.5 tech report · 2024
Cited alongside, same era.
Vidur: A large-scale simulation framework for llm inference, 2024
Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav Gulavani, Ramachandran Ramjee, and Alexey Tumanov · 2024
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve, 2024
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee · 2024
Cited alongside, same era.
Shortcut-connected expert parallelism for accelerating mixture-of-experts, 2024
Weilin Cai, Juyong Jiang, Le Qin, Junwei Cui, Sunghun Kim, and Jiayi Huang · 2024
Cited alongside, same era.
Luke Merrick, Danmei Xu, Gaurav Nuti, and Daniel Campos · 2024
Closest in time.
Splitwise: Efficient generative llm inference using phase splitting, 2024
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Closest in time.
Mooncake: A kvcache-centric disaggregated architecture for llm serving, 2024
Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu · 2024
Closest in time.
Scattered mixture-of-experts implementation, 2024
Shawn Tan, Yikang Shen, Rameswar Panda, and Aaron Courville · 2024
Closest in time.
Hmoe: Heterogeneous mixture of experts for language modeling, 2024
An Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu, Zhen Yang, Pinxue Zhao, J. N. Han, Zhanhui Kang, Di Wang, Naoaki Okazaki, and Cheng zhong Xu · 2024
Closest in time.
Llm inference unveiled: Survey and roofline model insights, 2024
Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving, 2024
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Nanoflow: Towards optimal large language model serving throughput, 2024
Kan Zhu, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Yufei Gao, Qinyu Xu, Tian Tang, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, Stephanie Wang, Arvind Krishnamurthy, and Baris Kasikci · 2024
Closest in time.