Fetching the paper…
Reading the bibliography…
Mixture of Experts (MoE) LLMs have recently gained attention for their ability to enhance performance by selectively engaging specialized subnetworks or "experts" for each input.
A study of replacement algorithms for a virtual-storage computer
Laszlo A. Belady · 1966
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Task-specific expert pruning for sparse mixture-of-experts
Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Optimizing dynamic neural networks with brainstorm
Weihao Cui, Zhenhua Han, Lingji Ouyang, Yichuan Wang, Ningxin Zheng, Lingxiao Ma, Yuqing Yang, Fan Yang, Jilong Xue, Lili Qiu, Lidong Zhou, Quan Chen, Haisheng Tan, and Minyi Guo · 2023
Earlier work this paper cites.
Fast inference of mixture-of-experts language models with offloading
Artyom Eliseev and Denis Mazur · 2023
Earlier work this paper cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Earlier work this paper cites.
Optq: Accurate quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2023
Earlier work this paper cites.
Towards moe deployment: Mitigating inefficiencies in mixture-of-expert (moe) inference
Haiyang Huang, Newsha Ardalani, Anna Sun, Liu Ke, Hsien-Hsin S Lee, Anjali Sridhar, Shruti Bhosale, Carole-Jean Wu, and Benjamin Lee · 2023
Earlier work this paper cites.
Deepum: Tensor migration and prefetching in unified memory
Jaehoon Jung, Jinpyo Kim, and Jaejin Lee · 2023
Earlier work this paper cites.
Swapmoe: Serving off-the-shelf moe-based large language models with tunable memory budget
Rui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Xiaozhou Ye, Ye Ouyang, Linghe Kong, and Yunxin Liu · 2023
Earlier work this paper cites.
Edgemoe: Fast on-device inference of moe-based large language models
Rongjie Yi, Liwei Guo, Shiyun Wei, Ao Zhou, Shangguang Wang, and Mengwei Xu · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, and Misha Bilenko et al · 2024
Cited alongside, same era.
LLM in a flash: Efficient large language model inference with limited memory
Keivan Alizadeh, Seyed-Iman Mirzadeh, Dmitry Belenko, S. Khatamifard, Minsik Cho, Carlo C. del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar · 2024
Cited alongside, same era.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, and Hao Yang et al · 2024
Cited alongside, same era.
Olmoe: Open mixture-of-experts language models
Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi, Noah A. Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, and Shyamal Anadkat et al · 2024
Closest in time.
Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters, February 2024
Qwen Qwen Team · 2024
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models
Zhixu Du, Shiyu Li, Yuhao Wu, Xiangyu Jiang, Jingwei Sun, Qilin Zheng, Yongkai Wu, Ang Li, Hai Li, and Yiran Chen · 2024
Cited alongside, same era.
llama.cpp, 2023
Georgi Gerganov · 2024
Cited alongside, same era.
Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference
Ranggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang, Xiaohu Tang, Ting Cao, and Mao Yang · 2024
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Cited alongside, same era.
Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models
Keisuke Kamahori, Yile Gu, Kan Zhu, and Baris Kasikci · 2024
Cited alongside, same era.
Merge, then compress: Demystify efficient SMoe with hints from its routing policy
Pingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung, Yu Cheng, Mohit Bansal, and Tianlong Chen · 2024
Cited alongside, same era.
Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models
Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, and Hongsheng Li · 2024
Cited alongside, same era.
Scaling laws for fine-grained mixture of experts
Jan Ludziejewski, Jakub Krajewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, Krystian Król, Tomasz Odrzygóźdź, Piotr Sankowski, et al · 2024
Cited alongside, same era.
Peng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu, Jing Wang, Pheng-Ann Heng, Chao Li, and Minyi Guo · 2024
Closest in time.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, and Andrew M. Dai et al · 2024
Closest in time.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Closest in time.
Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services
Dianhai Yu, Liang Shen, Hongxiang Hao, Weibao Gong, Huachao Wu, Jiang Bian, Lirong Dai, and Haoyi Xiong · 2024
Closest in time.
Adapmoe: Adaptive sensitivity-based expert gating and management for efficient moe inference
Shuzhang Zhong, Ling Liang, Yuan Wang, Runsheng Wang, Ru Huang, and Meng Li · 2024
Closest in time.
Efficient inference offloading for mixture-of-experts large language models in internet of medical things
Xiaoming Yuan, Weixuan Kong, Zhenyu Luo, and Minrui Xu · 2077
Closest in time.
Efficient inference offloading for mixture-of-experts large language models in internet of medical things
Xiaoming Yuan, Weixuan Kong, Zhenyu Luo, and Minrui Xu · 2079
Closest in time.