Fetching the paper…
Reading the bibliography…
In large language models like the Generative Pre-trained Transformer, the Mixture of Experts paradigm has emerged as a powerful technique for enhancing model expressiveness and accuracy.
S. Masoudnia and R. Ebrahimpour, “Mixture of experts: a literature survey,” Artificial Intelligence Review , vol. 42, pp. 275–293, 2014
2014
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. I. for AI, “C4: The colossal clean crawled corpus,” https://huggingface.co/datasets/allenai/c4 , 2020
2020
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pinto, D. Keysers, and N. Houlsby, “Scaling Vision with Sparse Mixture of Experts,” 2021
2021
Cited alongside, same era.
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, and A. Peste, “Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,” The Journal of Machine Learning Research , vol. 22, no. 1, pp. 10 882–11 005, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
S. Shen, Z. Yao, C. Li, T. Darrell, K. Keutzer, and Y. He, “Scaling Vision-Language Models with Sparse Mixture of Experts,” 2023
2023
Later among the works it cites.
OpenAI, “GPT-4 Technical Report,” 2023
2023
Later among the works it cites.
J. Li, Y. Jiang, Y. Zhu, C. Wang, and H. Xu, “Accelerating Distributed MoE Training and Inference with Lina,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23) , 2023, pp. 945–959
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nvidia, “Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, the World’s Largest and Most Powerful Generative Language Model,” 2021. [Online]. Available: https://developer.nvidia.com/blog/
2021
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” The Journal of Machine Learning Research , vol. 23, no. 1, pp. 5232–5270, 2022
2022
Cited alongside, same era.
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon et al. , “Mixture-of-experts with expert choice routing,” Advances in Neural Information Processing Systems , vol. 35, pp. 7103–7114, 2022
2022
Cited alongside, same era.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” in International Conference on Machine Learning . PMLR, 2022, pp. 18 332–18 346
2022
Cited alongside, same era.
J. He, J. Zhai, T. Antunes, H. Wang, F. Luo, S. Shi, and Q. Li, “FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models,” in Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2022, pp. 120–134
2022
Cited alongside, same era.
C. Chen, M. Li, Z. Wu, D. Yu, and C. Yang, “TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training,” Advances in Neural Information Processing Systems , vol. 35, pp. 22 173–22 186, 2022
2022
Cited alongside, same era.
NVIDIA, “Nvlink,” 2022. [Online]. Available: https://www.nvidia.com/en-us/data-center/nvlink/
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
Microsoft, “Megatron-deepspeed,” 2023. [Online]. Available: https://github.com/microsoft/Megatron-DeepSpeed
2023
Later among the works it cites.
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. H. Jha, S. Kumar, L. Lucy, X. Lyu, I. Magnusson, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. E. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, E. P. Walsh, H. Hajishirzi, N. A. Smith, L. Zettlemoyer, I. Beltagy, D. Groeneveld, J. Dodge, and K. Lo, “Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research,” arXiv preprint , 2023
2023
Later among the works it cites.
S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y. He, and A. Bhatele, “A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training,” in Proceedings of the 37th International Conference on Supercomputing , ser. ICS ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 203–214. [Online]. Available: https://doi.org/10.1145/3577193.3593704
2023
Later among the works it cites.
C. Hwang, W. Cui, Y. Xiong, Z. Yang, Z. Liu, H. Hu, Z. Wang, R. Salas, J. Jose, P. Ram et al. , “Tutel: Adaptive mixture-of-experts at scale,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Later among the works it cites.
M. Zhai, J. He, Z. Ma, Z. Zong, R. Zhang, and J. Zhai, “SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23) , 2023, pp. 961–975
2023
Later among the works it cites.
“Yelp dataset,” https://www.yelp.com/dataset , accessed: 2024-01-15
2024
Closest in time.