Fetching the paper…
Reading the bibliography…
MoE facilitates the development of large models by making the computational complexity of the model no longer scale linearly with increasing parameters.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation , vol. 3, no. 1, pp. 79–87, 1991
1991
Earlier work this paper cites.
2017
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby, “Scaling vision with sparse mixture of experts,” Advances in Neural Information Processing Systems , vol. 34, pp. 8583–8595, 2021
2021
Earlier work this paper cites.
S. Roller, S. Sukhbaatar, J. Weston et al. , “Hash layers for large sparse models,” Advances in Neural Information Processing Systems , vol. 34, pp. 17 555–17 566, 2021
2021
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022
2022
Cited alongside, same era.
B. Mustafa, C. Riquelme, J. Puigcerver, R. Jenatton, and N. Houlsby, “Multimodal contrastive learning with limoe: the language-image mixture of experts,” Advances in Neural Information Processing Systems , vol. 35, pp. 9564–9576, 2022
2022
Cited alongside, same era.
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon et al. , “Mixture-of-experts with expert choice routing,” Advances in Neural Information Processing Systems , vol. 35, pp. 7103–7114, 2022
2022
Cited alongside, same era.
X. Nie, X. Miao, Z. Wang, Z. Yang, J. Xue, L. Ma, G. Cao, and B. Cui, “Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,” Proceedings of the ACM on Management of Data , vol. 1, no. 1, pp. 1–19, 2023
2023
Later among the works it cites.
W. Wang, Z. Lai, S. Li, W. Liu, K. Ge, Y. Liu, A. Shen, and D. Li, “Prophet: Fine-grained load balancing for parallel training of large-scale moe models,” in 2023 IEEE International Conference on Cluster Computing (CLUSTER) . IEEE, 2023, pp. 82–94
2023
Later among the works it cites.
Y. Zhou, N. Du, Y. Huang, D. Peng, C. Lan, D. Huang, S. Shakeri, D. So, A. M. Dai, Y. Lu et al. , “Brainformers: Trading simplicity for efficiency,” in International Conference on Machine Learning . PMLR, 2023, pp. 42 531–42 542
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…