Fetching the paper…
Reading the bibliography…
Mixture of Experts (MoE) has become a key ingredient for scaling large foundation models while keeping inference costs steady.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Cited alongside, same era.
Scaling vision with sparse mixture of experts, 2021
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pinto, D. Keysers, and N. Houlsby · 2021
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts, 2022
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
W. Fedus, B. Zoph, and N. Shazeer · 2022
Later among the works it cites.
Megablocks: Efficient sparse training with mixture-of-experts
T. Gale, D. Narayanan, C. Young, and M. Zaharia · 2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…