Fetching the paper…
Reading the bibliography…
The increasing demand for deploying large Mixture-of-Experts (MoE) models in resource-constrained environments necessitates efficient approaches to address their high memory and computational requirements challenges.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
The penn treebank: Annotating predicate argument structure
Mitch Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger · 1994
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Improving the accuracy and hardware efficiency of neural networks using approximate multipliers
Mohammad Saeed Ansari, Vojtech Mrazek, Bruce F Cockburn, Lukas Sekanina, Zdenek Vasicek, and Jie Han · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Scalable and efficient moe training for multitask multilingual models
Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla · 2021
Earlier work this paper cites.
Taming sparsely activated transformer with stochastic experts
Simiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim, Hany Hassan, Ruofei Zhang, Tuo Zhao, and Jianfeng Gao · 2021
Cited alongside, same era.
M6-10t: A sharing-delinking paradigm for efficient multi-trillion parameter pretraining
Junyang Lin, An Yang, Jinze Bai, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Yong Li, Wei Lin, et al · 2021
Cited alongside, same era.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production
Young Jin Kim, Rawn Henry, Raffy Fahim, and Hany Hassan Awadalla · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Llm-qat: Data-free quantization aware training for large language models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
The case for 4-bit precision: k-bit inference scaling laws
Tim Dettmers and Luke Zettlemoyer · 2023
Later among the works it cites.
Moe-infinity: Activation-aware expert offloading for efficient moe serving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Cited alongside, same era.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Cited alongside, same era.
Fast inference of mixture-of-experts language models with offloading
Artyom Eliseev and Denis Mazur · 2023
Cited alongside, same era.
Qmoe: Practical sub-1-bit compression of trillion-parameter models
Elias Frantar and Dan Alistarh · 2023
Cited alongside, same era.
Mixture of quantized experts (moqe): Complementary effect of low-bit quantization and robustness
Young Jin Kim, Raffy Fahim, and Hany Hassan Awadalla · 2023
Cited alongside, same era.
Sparse backpropagation for moe training
Liyuan Liu, Jianfeng Gao, and Weizhu Chen · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang · 2023
Cited alongside, same era.
Leyang Xue, Yao Fu, Zhan Lu, Luo Mai, and Mahesh Marina · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models
Keisuke Kamahori, Yile Gu, Kan Zhu, and Baris Kasikci · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2024
Closest in time.
Billm: Pushing the limit of post-training quantization for llms
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, and Xiaojuan Qi · 2024
Closest in time.
Demystifying the compression of mixture-of-experts through a unified framework
Shwai He, Daize Dong, Liang Ding, and Ang Li · 2024
Closest in time.