Fetching the paper…
Reading the bibliography…
Sparsely activated Mixture-of-Experts (SMoE) has shown promise in scaling up the learning capacity of neural networks.
Importance estimation for neural network pruning, 2019
Molchanov, P., Mallya, A., Tyree, S., Frosio, I., and Kautz, J · 1906
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks, 2015
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming, 2017
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference, 2017
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2017
Earlier work this paper cites.
Coreset-based neural network compression, 2018
Dubey, A., Chatterjee, M., and Ahuja, N · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Earlier work this paper cites.
Towards a deep and unified understanding of deep neural models in nlp
Guan, C., Wang, X., Zhang, Q., Chen, R., He, D., and Xie, X · 2019
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P · 2019
Earlier work this paper cites.
Nvidia a100 tensor core gpu architecture
Nvidia · 2020
Earlier work this paper cites.
Stable rank normalization for improved generalization in neural networks and gans
Sanyal, A., Torr, P. H., and Dokania, P. K · 2020
Earlier work this paper cites.
St-moe: Designing stable and transferable sparse expert models
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2020
Earlier work this paper cites.
Knowledge distillation: A survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D · 2021
Earlier work this paper cites.
Scalable and efficient moe training for multitask multilingual models, 2021
Kim, Y. J., Awan, A. A., Muzio, A., Salinas, A. F. C., Lu, L., Hendy, A., Rajbhandari, S., He, Y., and Awadalla, H. H · 2021
Earlier work this paper cites.
Accelerating sparse deep neural networks, 2021
Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P · 2021
Cited alongside, same era.
Learning n: m fine-grained structured sparse neural networks from scratch
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts, 2022
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., Anantharaman, G., Li, X., Chen, S., Akin, H., Baines, M., Martin, L., Zhou, X., Koura, P. S., O’Horo, B., Wang, J., Zettlemoyer, L., Diab, M., Kozareva, Z., and Stoyanov, V · 2022
Cited alongside, same era.
Task-specific expert pruning for sparse mixture-of-experts, 2022
Chen, T., Huang, S., Xie, Y., Jiao, B., Jiang, D., Zhou, H., Li, J., and Wei, F · 2022
Cited alongside, same era.
Koishekenov, Y., Berard, A., and Nikoulina, V · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Later among the works it cites.
Is c4 dataset optimal for pruning? an investigation of calibration data for llm pruning
Bandari, A., Yin, L., Hsieh, C.-Y., Jaiswal, A., Chen, T., Shen, L., Krishna, R., and Liu, S · 2024
Later among the works it cites.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Training your sparse neural network better with any mask
Jaiswal, A. K., Ma, H., Chen, T., Ding, Y., and Wang, Z · 2022
Cited alongside, same era.
Is a modular architecture enough?
Mittal, S., Bengio, Y., and Lajoie, G · 2022
Cited alongside, same era.
Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask?, 2022
Paul, M., Chen, F., Larsen, B. W., Frankle, J., Ganguli, S., and Dziugaite, G. K · 2022
Cited alongside, same era.
Rajbhandari, S., Li, C., Yao, Z., Zhang, M., Aminabadi, R. Y., Awan, A. A., Rasley, J., and He, Y · 2022
Cited alongside, same era.
Structural pruning via latency-saliency knapsack, 2022
Shen, M., Yin, H., Molchanov, P., Mao, L., Liu, J., and Alvarez, J. M · 2022
Cited alongside, same era.
Dai, D., Deng, C., Zhao, C., Xu, R. X., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., Xie, Z., Li, Y. K., Huang, P., Luo, F., Ruan, C., Sui, Z., and Liang, W · 2024
Later among the works it cites.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI, Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., Guo, D., Yang, D., Chen, D., Ji, D., Li, E., Lin, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Yang, H., Zhang, H., Ding, H., Xin, H., Gao, H., Li, H., Qu, H., Cai, J. L., Liang, J., Guo, J., Ni, J., Li, J., Chen, J., Yuan, J., Qiu, J., Song, J., Dong, K., Gao, K., Guan, K., Wang, L., Zhang, L., Xu, L., Xia, L., Zhao, L., Zhang, L., Li, M., Wang, M., Zhang, M., Zhang, M., Tang, M., Li, M., Tian, N., Huang, P., Wang, P., Zhang, P., Zhu, Q., Chen, Q., Du, Q., Chen, R. J., Jin, R. L., Ge, R., Pan, R., Xu, R., Chen, R., Li, S. S., Lu, S., Zhou, S., Chen, S., Wu, S., Ye, S., Ma, S., Wang, S., Zhou, S., Yu, S., Zhou, S., Zheng, S., Wang, T., Pei, T., Yuan, T., Sun, T., Xiao, W. L., Zeng, W., An, W., Liu, W., Liang, W., Gao, W., Zhang, W., Li, X. Q., Jin, X., Wang, X., Bi, X., Liu, X., Wang, X., Shen, X., Chen, X., Chen, X., Nie, X., Sun, X., Wang, X., Liu, X., Xie, X., Yu, X., Song, X., Zhou, X., Yang, X., Lu, X., Su, X., Wu, Y., Li, Y. K., Wei, Y. X., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zhao, Y., Sun, Y., Li, Y., Wang, Y., Zheng, Y., Zhang, Y., Xiong, Y., Zhao, Y., He, Y., Tang, Y., Piao, Y., Dong, Y., Tan, Y., Liu, Y., Wang, Y., Guo, Y., Zhu, Y., Wang, Y., Zou, Y., Zha, Y., Ma, Y., Yan, Y., You, Y., Liu, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Huang, Z., Zhang, Z., Xie, Z., Hao, Z., Shao, Z., Wen, Z., Xu, Z., Zhang, Z., Li, Z., Wang, Z., Gu, Z., Li, Z., and Xie, Z · 2024
Later among the works it cites.
Demystifying the compression of mixture-of-experts through a unified framework
He, S., Dong, D., Ding, L., and Li, A · 2024
Later among the works it cites.
Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression
Hong, J., Duan, J., Zhang, C., Li, Z., Xie, C., Lieberman, K., Diffenderfer, J., Bartoldson, B., Jaiswal, A., Xu, K., et al · 2024
Later among the works it cites.
From galore to welore: How low-rank weights non-uniformly emerge from low-rank gradients
Jaiswal, A., Yin, L., Zhang, Z., Liu, S., Zhao, J., Tian, Y., and Wang, Z · 2024
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Later among the works it cites.
Mlp can be a good transformer learner
Lin, S., Lyu, P., Liu, D., Tang, T., Liang, X., Song, A., and Chang, X · 2024
Later among the works it cites.
Lu, X., Liu, Q., Xu, Y., Zhou, A., Huang, S., Zhang, B., Yan, J., and Li, H · 2024
Later among the works it cites.
Seer-moe: Sparse expert efficiency through regularization for mixture-of-experts
Muzio, A., Sun, A., and He, C · 2024
Later among the works it cites.
Revisiting smoe language models by evaluating inefficiencies with task specific expert pruning, 2024
Sarkar, S., Lausen, L., Cevher, V., Zha, S., Brox, T., and Karypis, G · 2024
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models, 2024
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2024
Later among the works it cites.
Q-galore: Quantized galore with int4 projection and layer-adaptive low-rank gradients
Zhang, Z., Jaiswal, A., Yin, L., Liu, S., Zhao, J., Tian, Y., and Wang, Z · 2024
Later among the works it cites.
Galore: Memory-efficient llm training by gradient low-rank projection
Zhao, J., Zhang, Z., Chen, B., Wang, Z., Anandkumar, A., and Tian, Y · 2024
Later among the works it cites.
Idea prune: An integrated enlarge-and-prune pipeline in generative language model pretraining
Li, Y., Du, X., Jaiswal, A., Lei, T., Zhao, T., Wang, C., and Wang, J · 2025
Closest in time.