Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) architectures enable efficient scaling of large language models by activating only a subset of parameters per input.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; and Chen, Z. 2020 · 2006
Earlier work this paper cites.
A genetic algorithm using Calinski-Harabasz index for automatic clustering problem
Lima, S. P.; and Cruz, M. D. 2020 · 2020
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; de Oliveira Pinto, H. P.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; Ray, A.; Puri, R.; Krueger, G.; Petrov, M.; Khlaaf, H.; Sastry, G.; Mishkin, P.; Chan, B.; Gray, S.; Ryder, N.; Pavlov, M.; Power, A.; Kaiser, L.; Bavarian, M.; Winter, C.; Tillet, P.; Such, F. P.; Cummings, D.; Plappert, M.; Chantzis, F.; Barnes, E.; Herbert-Voss, A.; Guss, W. H.; Nichol, A.; Paino, A.; Tezak, N.; Tang, J.; Babuschkin, I.; Balaji, S.; Jain, S.; Saunders, W.; Hesse, C.; Carr, A. N.; Leike, J.; Achiam, J.; Misra, V.; Morikawa, E.; Radford, A.; Knight, M.; Brundage, M.; Murati, M.; Mayer, K.; Welinder, P.; McGrew, B.; Amodei, D.; McCandlish, S.; Sutskever, I.; and Zaremba, W. 2021 · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Towards Understanding Mixture of Experts in Deep Learning
Chen, Z.; Deng, Y.; Wu, Y.; Gu, Q.; and Li, Y. 2022 · 2022
Earlier work this paper cites.
Unified scaling laws for routed language models
Clark, A.; de Las Casas, D.; Guy, A.; Mensch, A.; Paganini, M.; Hoffmann, J.; Damoc, B.; Hechtman, B.; Cai, T.; Borgeaud, S.; et al. 2022 · 2022
Earlier work this paper cites.
Stablemoe: Stable routing strategy for mixture of experts
Dai, D.; Dong, L.; Ma, S.; Zheng, B.; Sui, Z.; Chang, B.; and Wei, F. 2022 · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W.; Zoph, B.; and Shazeer, N. 2022 · 2022
Cited alongside, same era.
MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Gale, T.; Narayanan, D.; Young, C.; and Zaharia, M. 2022 · 2022
Cited alongside, same era.
Patch-level routing in mixture-of-experts is provably sample-efficient for convolutional neural networks
Chowdhury, M. N. R.; Zhang, S.; Wang, M.; Liu, S.; and Chen, P.-Y. 2023 · 2023
Cited alongside, same era.
Teleqna: A benchmark dataset to assess large language models telecommunications knowledge
Maatouk, A.; Ayed, F.; Piovesan, N.; De Domenico, A.; Debbah, M.; and Luo, Z.-Q. 2023 · 2023
Cited alongside, same era.
Chen, F.; Li, P.; Hong, Z.; Su, Z.; and Guo, S. 2024 · 2024
Closest in time.
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Dai, D.; Deng, C.; Zhao, C.; Xu, R. X.; Gao, H.; Chen, D.; Li, J.; Zeng, W.; Yu, X.; Wu, Y.; Xie, Z.; Li, Y. K.; Huang, P.; Luo, F.; Ruan, C.; Sui, Z.; and Liang, W. 2024 · 2024
Closest in time.
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
Huang, Q.; An, Z.; Zhuang, N.; Tao, M.; Zhang, C.; Jin, Y.; Xu, K.; Xu, K.; Chen, L.; Huang, S.; and Feng, Y. 2024 · 2024
Closest in time.
Locmoe: A low-overhead moe for large language model training
Li, J.; Sun, Z.; He, X.; Zeng, L.; Lin, Y.; Li, E.; Zheng, B.; Zhao, R.; and Chen, X. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; and Bowman, S. R. 2023 · 2023
Cited alongside, same era.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Cited alongside, same era.
A Survey on Mixture of Experts in Large Language Models
Cai, W.; Jiang, J.; Wang, F.; Tang, J.; Kim, S.; and Huang, J. 2024 · 2024
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Hanna, E. B.; Bressand, F.; et al. 2024a
Cited in the paper.
{ \{ MegaScale } \} : Scaling Large Language Model Training to More Than 10,000 { \{ GPUs } \}
Jiang, Z.; Lin, H.; Zhong, Y.; Huang, Q.; Chen, Y.; Zhang, Z.; Peng, Y.; Li, X.; Xie, C.; Nong, S.; et al. 2024b
Cited in the paper.
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024a
Cited in the paper.
Routers in Vision Mixture of Experts: An Empirical Study
Liu, T.; Blondel, M.; Riquelme, C.; and Puigcerver, J. 2024b
Cited in the paper.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017a
Cited in the paper.
Pham, Q.; Do, G.; Nguyen, H.; Nguyen, T.; Liu, C.; Sartipi, M.; Nguyen, B. T.; Ramasamy, S.; Li, X.; Hoi, S.; and Ho, N. 2024 · 2024
Closest in time.
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
Xie, Z.; Zhang, Y.; Zhuang, C.; Shi, Q.; Liu, Z.; Gu, J.; and Zhang, G. 2024 · 2024
Closest in time.
Accelerating MoE Model Inference with Expert Sharding
Balmau, O.; Kermarrec, A.-M.; Pires, R.; Santo, A. L. E.; de Vos, M.; and Vujasinovic, M. 2025 · 2025
Closest in time.