Fetching the paper…
Reading the bibliography…
The Mixture-of-Experts (MoE) architecture has demonstrated significant advantages in the era of Large Language Models (LLMs), offering enhanced capabilities with reduced inference costs.
Adaptive mixture of local expert
Robert Jacobs, Michael Jordan, Steven Nowlan, and Geoffrey Hinton · 1991
Earlier work this paper cites.
The meta-pi network: building distributed knowledge representations for robust multisource pattern recognition
J.B. Hampshire and A. Waibel · 1992
Earlier work this paper cites.
Mixtures of gaussian processes
Volker Tresp · 2000
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Earlier work this paper cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He · 2021
Earlier work this paper cites.
Deepspeed- inference: Enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Switch transformers: scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Earlier work this paper cites.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan · 2022
Earlier work this paper cites.
Yukun Huang, Yanda Chen, Zhou Yu, and Kathleen McKeown · 2022
Earlier work this paper cites.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
Rishabh Agarwal, Nino Vieillard, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem · 2023
Earlier work this paper cites.
Fast inference of mixture-of-experts language models with offloading
Artyom Eliseev and Denis Mazur · 2023
Earlier work this paper cites.
ggerganov/llama.cpp: Port of facebook’s llama model in c/c++
Georgi Gerganov · 2023
Earlier work this paper cites.
Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2023
Earlier work this paper cites.
Beating google and apple, huawei brings large ai model to mobile voice assistant
Huawei · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Adaptive gating in mixture-of-experts based language models
Jiamin Li, Qiang Su, Yitao Yang, Yimin Jiang, Cong Wang, and Hong Xu · 2023
Cited alongside, same era.
Merge, then compress: Demystify efficient smoe with hints from its routing policy
Pingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung, Yu Cheng, Mohit Bansal, and Tianlong Chen · 2023
Cited alongside, same era.
Losparse: Structured compression of large language models based on low-rank and sparse approximation
Yixiao Li, Yifan Yu, Qingru Zhang, Chen Liang, Pengcheng He, Weizhu Chen, and Tuo Zhao · 2023
Cited alongside, same era.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI · 2024
Closest in time.
Hugging face hub
Hugging Face · 2024
Closest in time.
Llm-based edge intelligence: A comprehensive survey on architectures, applications, security and trustworthiness
Othmane Friha, Mohamed Amine Ferrag, Burak Kantarci, Burak Cakmak, Arda Ozgun, and Nassira Ghoualmi-Zine · 2024
Closest in time.
Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference
Ranggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang, Xiaohu Tang, Ting Cao, and Mao Yang · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2023
Cited alongside, same era.
World’s first on-device demonstration of stable diffusion on an android phone
Qualcomm AI Research · 2023
Cited alongside, same era.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen · 2023
Cited alongside, same era.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Bitnet: Scaling 1-bit transformers for large language models
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Cited alongside, same era.
Keisuke Kamahori, Yile Gu, Kan Zhu, and Baris Kasikci · 2024
Closest in time.
Shortened llama: A simple depth pruning for large language models
Bo-Kyeong Kim, Geonmin Kim, Tae-Ho Kim, Thibault Castells, Shinkook Choi, Junho Shin, and Hyoung-Kyu Song · 2024
Closest in time.
SwapMoE: Serving off-the-shelf MoE-based large language models with tunable memory budget
Rui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Xiaozhou Ye, Ye Ouyang, Linghe Kong, and Yunxin Liu · 2024
Closest in time.
Cats: Contextually-aware thresholding for sparsity in large language models
Je-Yong Lee, Donghyun Lee, Genghan Zhang, Mo Tiwari, and Azalia Mirhoseini · 2024
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen · 2024
Closest in time.
Geforce rtx 4090
NVIDIA · 2024
Closest in time.
Jetson agx orin developer kit
NVIDIA · 2024
Closest in time.
Samsung pcie 4.0 nvme ssd 980 pro
Samsung · 2024
Closest in time.
Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024
Qwen Team · 2024
Closest in time.
Svd-llm: Truncation-aware singular value decomposition for large language model compression
Xin Wang, Yu Zheng, Zhongwei Wan, and Mi Zhang · 2024
Closest in time.
Moe-infinity: Activation-aware expert offloading for efficient moe serving
Leyang Xue, Yao Fu, Zhan Lu, Luo Mai, and Mahesh Marina · 2024
Closest in time.
Powerinfer-2: Fast large language model inference on a smartphone
Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, and Haibo Chen · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu · 2024
Closest in time.
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
Shuzhang Zhong, Ling Liang, Yuan Wang, Runsheng Wang, Ru Huang, and Meng Li · 2024
Closest in time.