Efficient expert pruning for sparse mixture-of-experts language models: Enhancing performance and reducing inference costs
Original
Enshu Liu, Junyi Zhu, Zinan Lin, Xuefei Ning, Matthew B Blaschko, Shengen Yan, Guohao Dai, Huazhong Yang, and Yu Wang · 2024
Later among the works it cites.
Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models
Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, and Hongsheng Li · 2024
Later among the works it cites.
Seer-moe: Sparse expert efficiency through regularization for mixture-of-experts, 2024
Alexandre Muzio, Alex Sun, and Churan He · 2024
Later among the works it cites.
Libmoe: A library for comprehensive benchmarking mixture of experts in large language models
Original
Nam V Nguyen, Thong T Doan, Luong Tran, Van Nguyen, and Quang Pham · 2024
Later among the works it cites.
Revisiting smoe language models by evaluating inefficiencies with task specific expert pruning
Original
Soumajyoti Sarkar, Leonard Lausen, Volkan Cevher, Sheng Zha, Thomas Brox, and George Karypis · 2024
Later among the works it cites.
Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024
Qwen Team · 2024
Later among the works it cites.
Svd-llm: Truncation-aware singular value decomposition for large language model compression
Original
Xin Wang, Yu Zheng, Zhongwei Wan, and Mi Zhang · 2024
Later among the works it cites.
Moe-pruner: Pruning mixture-of-experts large language model using the hints from its router
Original
Yanyue Xie, Zhi Zhang, Ding Zhou, Cong Xie, Ziang Song, Xin Liu, Yanzhi Wang, Xue Lin, and An Xu · 2024
Later among the works it cites.
Moe-infinity: Activation-aware expert offloading for efficient moe serving, 2024
Leyang Xue, Yao Fu, Zhan Lu, Luo Mai, and Mahesh Marina · 2024
Later among the works it cites.
Qwen2.5 technical report
Original
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2024
Later among the works it cites.
Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decomposition
Original
Cheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang, Yu Gong, Yuanlin Duan, Wenqi Jia, Miao Yin, Yu Cheng, and Bo Yuan · 2024
Later among the works it cites.
HyperMoE: Towards better mixture of experts via transferring among experts
Hao Zhao, Zihan Qiu, Huijia Wu, Zili Wang, Zhaofeng He, and Jie Fu · 2024
Later among the works it cites.
Stbllm: Breaking the 1-bit barrier with structured binary llms
Peijie Dong, Lujun Li, Yuedong Zhong, Dayou Du, Ruibo Fan, Yuhan Chen, Zhenheng Tang, Qiang Wang, Wei Xue, Yike Guo, et al · 2025
Closest in time.
Delta decompression for moe-based LLMs compression
Hao Gu, Wei Li, Lujun Li, Zhu Qiyuan, Mark Lee, Shengjie Sun, Wei Xue, and Yike Guo · 2025
Closest in time.
Mixture compressor for mixture-of-experts LLMs gains more
Wei Huang, Yue Liao, Jianhui Liu, Ruifei He, Haoru Tan, Shiming Zhang, Hongsheng Li, Si Liu, and XIAOJUAN QI · 2025
Closest in time.
Structured mixture-of-experts LLMs compression via singular value decomposition
Wei Li, Lujun Li, You-Liang Huang, Mark G. Lee, Shengjie Sun, Wei Xue, and Yike Guo · 2025
Closest in time.